Ecommerce ops · Production

Shopify fine-tunes a tool-calling agent for Flow: 2.2x faster, 68% cheaper, outperforms closed models

The problem

Store owners who are not engineers found building automation workflows from a blank canvas in Shopify Flow daunting. The feature also faced a cold start problem: no production conversations existed to learn from because Sidekick had not yet been deployed.

First attempt

Offline benchmark results showed parity with the prompt-based agent, but initial production deployment revealed the fine-tuned model had a 35% lower workflow activation rate because synthetic training data did not cover real user requests such as editing existing workflows, handling email configurations, and working with third-party integrations.

Workflow diagram · grounded in source
1
Sample production workflows
Trigger
Thousands of anonymized store-owner workflows are sampled from production and filtered for quality to bootstrap training data.
source quote
“We reverse-engineered user intent from existing production workflows. Thousands of anonymized store owners had already built workflows manually in Flow. We sampled those and filtered for quality: workflows that had run at least once in the last seven days, from …”
2
Generate synthetic training examples
Ai action
A stronger LLM generates a plausible natural-language request for each sampled workflow, and the full multi-turn tool-call trajectory an ideal agent would execute is constructed.
source quote
“Generate a user query. Use a stronger LLM to produce a plausible natural-language request that would lead to this workflow. Construct the tool trajectory. Build the full multi-turn sequence of tool calls that an ideal agent would execute to arrive …”
3
Translate to Python DSL
Ai action
Switching from the native JSON DSL to Python for workflow representation improved syntactic correctness by 22 points and semantic correctness by 13 points.
source quote
“switching from the JSON DSL to the Python DSL improved syntactic correctness by 22 points and semantic correctness by 13 points”
4
Fine-tune Qwen3-32B
Ai action
Qwen3-32B is fine-tuned on the synthetic dataset.
source quote
“We fine-tuned Qwen3-32B on this synthetic dataset”
5
Evaluate against benchmark
Validation
An LLM evaluation framework checks semantic correctness against 300 hand-crafted examples and validates syntactic correctness programmatically.
source quote
“evaluated it against a benchmark of 300 hand-crafted examples covering the breadth of expected Flow usage. An LLM evaluation framework compares the generated workflow against the expected one for semantic correctness, and validates syntactic correctness programmatically”
6
Deploy to production traffic
Output
The model is deployed to a slice of production traffic to observe real-world performance.
source quote
“We deployed it to 1% of traffic to see how it held up”
7
Score, route, and retrain
Feedback loop
Every production conversation is scored by an LLM judge, and high-scoring examples are routed into the training pool automatically.
source quote
“Every production conversation becomes a training signal. We sample high-quality examples: conversations where merchants actually activated the workflow afterwards. The judge scores them, and high-scoring conversations are routed into the training pool automatically. Low-scoring ones are quarantined for review”
Reported outcome

The fine-tuned model is 2.2x faster and 68% cheaper than the closed-model baseline, outperforms closed models, and now serves the majority of production traffic, with a continuous weekly retraining flywheel that closes quality gaps identified in production.

Reported metrics
Inference speed improvement2.2x faster
Inference cost reduction68% cheaper
Workflow activation rate gap at initial deployment vs prompt-based agent35% lower
syntactic correctness improvement (Python DSL vs JSON DSL)22 points
Show all 7 reported metrics
inference speed improvement2.2x faster
inference cost reduction68% cheaper
workflow activation rate gap at initial deployment vs prompt-based agent35% lower
syntactic correctness improvement (Python DSL vs JSON DSL)22 points
semantic correctness improvement (Python DSL vs JSON DSL)13 points
share of failures from email workflows25%
share of failures from diverse condition patterns16%
Reported stack
Qwen3-32BH200 GPUsFSDPTangleCometMLHuggingFaceCentMLSidekick
◆ Does this fit your context?

Compare to your context

Tell us your scale, team, and constraints. We'll show what changes at your size, what fails at your scale, and whether this case is a fit, needs adaptation, or won't scale to you. Free demo, no signup.

Compare to your context →
~30 seconds · free
Source
https://shopify.engineering/fine-tuning-agent-shopify-flow
Read source ↗

Frequently asked questions

What did this team achieve with this AI workflow?

The fine-tuned model is 2.2x faster and 68% cheaper than the closed-model baseline, outperforms closed models, and now serves the majority of production traffic, with a continuous weekly retraining flywheel that close…

What tools did this team use?

Qwen3-32B, H200 GPUs, FSDP, Tangle, CometML, HuggingFace, CentML, Sidekick.

What results were reported?

Inference speed improvement: 2.2x faster; Inference cost reduction: 68% cheaper; Workflow activation rate gap at initial deployment vs prompt-based agent: 35% lower; syntactic correctness improvement (Python DSL vs JSON DSL): 22 points (source-reported, not independently verified).

What failed first in this deployment?

Offline benchmark results showed parity with the prompt-based agent, but initial production deployment revealed the fine-tuned model had a 35% lower workflow activation rate because synthetic training data did not cov…

How is this ecommerce ops AI workflow structured?

Sample production workflows → Generate synthetic training examples → Translate to Python DSL → Fine-tune Qwen3-32B → Evaluate against benchmark → Deploy to production traffic → Score, route, and retrain.

WHAT TO DO WITH THIS

Now compare it to your context

This case is one data point. Whether its pattern fits you depends on your volumes, your stack, and your exception load — that comparison is the step no case study can do for you.