ServiceNow's CoreAI team just published AutoSynthData, a pipeline for turning an agent's specific weaknesses into training data. It's one of the clearer explorations I've seen of the gap between "we can generate synthetic tasks" and "we can train on them safely." The architecture addresses a problem that's becoming urgent as teams move agents into production: a broadly capable model can still fail in your environment, on your workflows, with your constraints. You need training data that targets those gaps, not generic instruction-following.
What makes this interesting isn't the idea of synthetic data—everyone's doing that now—but the explicit quality-control loop and the way it uses evaluation failures to steer curriculum.
The Core Insight: Generation Is Easy, Validation Is Hard
The AutoSynthData framing starts with three properties a useful training task must satisfy:
- Feasibility: There exists at least one trajectory that solves the task given the current environment's tools and state
- Realism: The user prompt resembles something someone would actually request
- Difficulty: The task exposes a weakness the current model has but a stronger teacher can solve
They also specify what makes a verifier reliable: consistency with the task specification, soundness (rejecting bad solutions), and completeness (accepting valid solutions beyond a single reference trajectory).
This matters because a lax verifier rewards incorrect behavior during training, while an overly restrictive one penalizes valid approaches. You can't just generate plausible-looking tasks and hope the model learns the right thing.
From Failures to Curriculum
The pipeline starts by running both the target model and a stronger teacher on diagnostic tasks in the target environment. It identifies patterns: which capabilities are being tested, which tools and workflows are involved, where the target fails and how the teacher succeeds.
Those findings get distilled into what they call "capability specification cards." Critically, the cards are sanitized—the generator doesn't see original prompts, entities, trajectories, or verifier details from the evaluation set. It receives the abstract capability and creates new tasks with different prompts, states, and solution paths.
This separation keeps the synthetic data from overfitting to evaluation examples while ensuring it targets the same underlying weaknesses.
Two-Phase Generation: Target and Multiply
AutoSynthData builds datasets in two phases:
Target phase: Workers generate independent tasks from capability specifications in parallel. Each candidate goes through validation, execution, solver evaluation, and repair before acceptance. This creates a core set of vetted examples.
Multiply phase: The system expands the dataset by creating novel variants of accepted target samples. Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier. Crucially, a multiplied sample cannot seed another multiplied sample—this anchors expansion to the vetted target set and limits drift.
The architecture separates generation control from environment-specific execution through an adapter pattern. The controller coordinates generation and quality control; the adapter handles environment execution, task management, reference replay, and deterministic verification.
The Quality-Control Loop
This is where the work gets real. Every candidate must clear multiple gates before entering training data:
Solver Evaluation
The system measures difficulty by running both the target model and the stronger solver multiple times. In their configuration, they favor tasks the target solves on at most one of three trials and the solver solves on at least two of three. This filters out tasks that are too easy or unsolvably hard.
Positive and Negative Verification
Positive verification asks: does the intended solution actually solve the generated task? The pipeline executes the reference trajectory in the environment and checks the resulting state against the verifier. This catches mismatches between the prompt, initial state, solution, and success criteria.
Negative verification asks: do incorrect outcomes fail? The system can mutate parts of the expected outcome and confirm those states no longer pass. This catches weak verifiers that award success without requiring the intended behavior.
Critique and Repair
Failed candidates don't get immediately discarded. A critic examines the sample and its failure mode, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or mismatches with the intended capability.
The diagnosis guides targeted repairs with a fixed retry limit. A repaired task must pass the gates again. This bounded repair process recovers salvageable candidates without infinite retry loops.
Batch-Level Review
Individually valid samples can still form a repetitive or unbalanced dataset. AutoSynthData includes a meta-review that examines accepted samples, rejected samples, and generation behavior across each batch.
It asks:
- Which task families are overrepresented, and which capability dimensions are missing?
- Are the same examples appearing repeatedly?
- Do particular targets keep failing generation?
- Are systematic problems appearing in critiques?
The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs work toward gaps. This prevents the dataset from collapsing into a few easy task families.
EnterpriseOps Gym Results
ServiceNow demonstrates the pipeline with their EnterpriseOps Gym environment. They show results on two domains: a "Hybrid" environment and an "ITSM" (IT Service Management) environment.
The pattern holds across both: models trained on AutoSynthData-generated curricula improve on the capabilities the curriculum targeted. As the model improves, the curriculum shifts toward what it still finds difficult.
They're explicit about this being a closed-loop process: evaluate the updated model, identify remaining gaps, generate the next round of training data.
Why This Matters
Most synthetic data work focuses on generation: better prompts, more diverse sampling, different model sizes. AutoSynthData makes the case that verification and curriculum design matter as much as generation quality.
The verification loops aren't nice-to-have polish—they're load-bearing. Without positive and negative verification, you're training on tasks that may be unsolvable, incorrectly solved, or rewarding the wrong behavior. Without batch-level review, your dataset collapses into repetitive examples.
The curriculum aspect is equally important. Generic instruction-following data won't fix specific capability gaps in a particular environment. You need training data that targets the model's actual weaknesses and scales up variation around those weaknesses.
What's Missing
The post focuses on the pipeline architecture and doesn't share:
- Specific performance numbers or benchmark comparisons
- Details about the stronger teacher model used
- Dataset sizes or generation costs
- How much human involvement the critique and repair steps require
- How the capability specification cards are created (automated extraction vs. human authoring)
These gaps make it hard to assess whether AutoSynthData is practical for teams without ServiceNow's resources. The verification loops sound expensive—multiple solver runs per candidate, execution in the target environment, critique and repair cycles.
But the architectural choices are sound. If you're building training data for agents in specific environments, you need some version of this: curriculum driven by failures, generation anchored to capability gaps, verification that candidates are solvable and correctly labeled, and batch-level quality control.
The Bigger Pattern
AutoSynthData sits in a broader shift: synthetic data is moving from "generate more examples" to "generate the right examples with verified labels." We're seeing similar patterns in other domains—Constitutional AI's critique loops, STaR's self-improvement cycles, various forms of verified synthetic reasoning data.
The common thread is that generation alone isn't enough. You need verification, filtering, and curriculum design that responds to what the model actually struggles with.
ServiceNow's contribution is making this concrete for enterprise agent training and showing what the verification infrastructure looks like when you take it seriously. The gap between "we generated 100k synthetic tasks" and "we generated 100k validated, non-redundant tasks targeting capability gaps" turns out to be most of the work.