Migration guide
A CSV is a snapshot of one moment, in one implicit market, with every value flattened to a string. It has served as the interchange format for product data for thirty years because it asks nothing of either end.
That is also its limit. There is no column for which market, no column for when this was true, and no way to express a value that is a structured thing rather than text. Catalogs, metafields and metaobjects exist to hold exactly those.
This is the mapping — where each column goes, in what order to move them, and what breaks if you skip a step.
Companion to REST is not GraphQL — ten places the missing dimension shows up, which argues why this matters. This one is the how.
Open a typical export and the columns fall into six groups, even though the file treats them identically. Naming them is most of the migration:
| CSV column | What it really is | Where it belongs |
|---|---|---|
| Handle, Title, Body, Vendor, Type, Tags | Identity and copy for the product | Product |
| Option1/2/3, SKU, Barcode, Weight, Inventory | The sellable unit | Variant |
| Price, Compare-at price | An offer, in one currency, for one market | Catalog price list, not a product field |
| Custom columns — care, spec, warranty, origin | Structured attributes flattened to text | Metafields, or metaobjects when reused |
| Duplicate rows per language | The same product, described differently | Translations |
| Duplicate rows per country, or a Published flag | Where the product is offered | Markets and publications |
The two rows that matter most are price and the duplicates. Everything else is a rename. Those two are where the file stops being a file and starts being a dimension.
The distinction trips people because a CSV has no equivalent of either. The test is whether the value is owned by one product or referenced by many.
A value that belongs to this product and nothing else. Care instructions, a spec sheet URL, a compliance note. It lives on the product, it is typed, and nothing else points at it.
A thing in its own right that products refer to. A material, a certification, a size chart, a manufacturer. Define it once, reference it from a hundred products, edit it in one place.
The CSV symptom of a missing metaobject is a column repeated identically across thousands of rows — "Recycled polyester (GRS certified)" written out ten thousand times. That value wanted to be an object with a name, a certificate number, and an expiry, and the file could only hold the label.
Order matters. Metaobject definitions are schema. They must exist before anything can reference them, and the fields on a definition are typed — changing a type after entries exist is disruptive. Design the definitions first, import second.
This is the conceptual break. In the CSV, price is a property of the product. In a catalog model, price is a property of a context — a market, a B2B location, or a sales channel — and the same product carries different prices in different contexts without being duplicated.
Practically that means the migration is not a column rename. One CSV row with one price becomes one product plus n price entries across n catalogs. If you previously handled a second country by exporting a second file with different numbers, those two files collapse into one product with two contexts.
The payoff is that the second market stops being another copy to reconcile. Adding a country becomes adding a context to an existing record — which is the difference between spend scaling with markets and spend scaling with copies.
Most large CSVs carry the same product several times: once per language, sometimes once per country. Those are not different products, and importing them as such is the single most expensive mistake in this migration — it puts duplication into the new model, where it will be harder to unwind than it was in the file.
Language duplicates become translations of one product. Country duplicates become market availability plus a catalog price. Keep them separate in your head: language is not market. German is spoken in three markets at three price points, and Belgium is one market with two languages. A model that fuses them will produce the wrong currency on a correct locale, and hreflang that annotates the wrong axis.
Sequence matters more than tooling. Each step depends on the one before it:
1. Classify every column into the six groups above. Do this on the real export, not a sample — the odd columns at the far right are usually the interesting ones.
2. Design and create metaobject definitions. Schema first. Nothing can reference a definition that does not exist, and retyping later is painful.
3. Import products and variants without prices. Get identity and structure correct while it is still cheap to fix.
4. Create markets and catalogs, then attach price lists per context. This is where the second country stops being a second file.
5. Attach metafields and metaobject references.
6. Add translations against the single product, replacing the language duplicates.
7. Read it back with a bulk operation to verify. One job returns the whole catalogue as a result file, which is both the verification step and the proof the shape is now queryable in one pass.
Once the dimensions exist, you stop pre-flattening. A query describes the shape you want — product, variants, the metaobjects it references, the price in this market, the translation for this locale — and the response arrives already resolved for that context.
That is the whole difference. The CSV workflow required you to decide the shape in advance and generate one file per audience. The catalog model resolves the shape at request time, from one record, and the file becomes an export format rather than the source of truth.
Worth saying plainly: moving to GraphQL does not, by itself, get you any of this. A shaped query over a flat catalogue returns flat data faster. The dimensions have to exist in the model — that is the migration. Transport is the easy half.
Round-tripping through CSV. Exporting the migrated catalogue back to a file drops the dimensions again. Once price is contextual, no single flat export is correct — it can only be correct for one market. Treat exports as views, not backups.
Typing regret on metaobjects. A field defined as text because the CSV had text is hard to convert to a reference or a date after ten thousand entries exist. Spend the time on definitions.
Country-to-language mapping. Covered above and still the most common defect in production.
Assuming history came with it. The migration moves current state. It does not create the price history you did not previously keep — that only begins accruing from the day you start recording it.
PIM Sync does this migration against a live Shopify catalogue and then keeps the record it produces — market-resolved, appended rather than overwritten, timestamped — beside whatever product master and feed manager you already run.
PIM Sync on the Shopify App Store ↗ · Read the argument behind it ↗