When the model you need doesn't exist
Last week, a HuggingFace engineer wanted a smaller version of the prompt rewriter that ships with Qwen-Image 2.1. The official model is 9B parameters, needs 20 GB of memory, and burns thousands of tokens before writing a single paragraph. The solution? Describe the requirement to ML Intern, HuggingFace's agentic workflow tool, and wait a day. The result: a 0.8B model that runs on CPU, returns valid output 99.7% of the time, uses a quarter of the teacher's tokens, and cost USD 16 in compute to build.
Then the engineer built six more models the same way. Each started as a conversational prompt in HuggingChat with ML Intern mode enabled. Each ended as a public model on the Hub with evaluation results in the model card. ML Intern planned the work, asked for budget approval before spending, ran smoke tests before full training runs, then trained, evaluated, and published—all on HuggingFace infrastructure.
This is what agentic workflows look like when they actually ship production artifacts.
The prompt is the spec
The initial prompts ran 450 words. By the sixth project, they'd grown to 2,000 words, because each project surfaced requirements for the next. All seven prompts are on GitHub, exactly as written.
A good prompt starts with the goal in one line and the motivation. Then it names exact components: dataset, base model, training script. Anything already verified goes under a literal "Verified facts, do not re-derive" heading, so the agent spends cycles on work instead of rediscovering known constraints. For a camera-angle LoRA, that section listed which trainer had just added transparent-image support and which open GitHub issues made the fallback trainer risky.
Two lines matter most. The first requests a baseline before training: "Report the base model's zero-shot score on the same metric before training so we can see the gain." Without it, you get a trained model and no reference point. The second is a smoke test with a checkpoint: for image LoRAs, the prompt asked for 50 training steps, then verification that saved weights had actually changed, before paying for the full run.
The prompt closes with deliverables and a spending cap: "Cap total spend at USD 12 and ask me before exceeding it." ML Intern starts every task with zero budget and requires permission before executing paid jobs, so the limit holds. When you omit a budget, the agent suggests paths based on project scope and asks which you prefer.
What got built
A citrus disease classifier
General vision models can describe a yellowing citrus leaf. Diagnosing whether it's a mite problem or magnesium deficiency, then recommending bio and non-bio treatments, is harder. The engineer assembled a training dataset from three sources hosted by Project-AgML: 3,017 annotated images across 21 pests, diseases, nutritional deficiencies, and treatments.
ML Intern fine-tuned Qwen3.5-2B on the dataset after benchmarking the foundation model first. On 335 test photos, the base model identified the correct problem 14.9% of the time. After two epochs on one A10G, the fine-tuned model hit 52.8%. Compute cost: USD 1.90.
A character LoRA for brand assets
Image models know plenty of characters. Huggy, drawn in the flat style of HuggingFace brand assets, was not one of them. ML Intern trained a LoRA on FLUX.2 klein base 4B using 84 captioned drawings.
The agent saved checkpoints every 100 steps and generated the same prompt set with each, making selection straightforward. Step 200 was the first where Huggy appeared on-model. From step 500 onward, Huggy's style started bleeding into unrelated prompts. The trained LoRA also works on the distilled klein model at 4 steps. Compute cost: USD 7.60.
Camera-angle LoRA for Qwen-Image 2.1
Camera-angle LoRAs rank among the most-liked community add-ons for earlier Qwen-Image models. You show the model an object and ask to see it 45 degrees from the left. Days after the Qwen-Image 2.1 release, nobody had built one.
ML Intern rendered 1,030 scanned household objects from Google Scanned Objects at 24 angles each—24,722 transparent images—on a CPU job costing cents. It selected 461 objects for training, held out 40 for testing, and generated 1,844 before-and-after pairs across 23 camera instructions.
Training ran 2,000 steps in 90 minutes on one A100 (USD 3.75). The full project took half a day and 48 jobs, counting failures from missing packages or wrong paths that required resubmission. Total compute: USD 16.
Doodle-in LoRA for object insertion
Upload a photo with a magenta scribble. Add a short prompt naming an object. The LoRA replaces the scribble with that object while preserving original lighting and composition.
No dataset existed, so the prompt described how to build one: start from a real photo in Open Images, remove one object with the LaMa inpainting model, draw a scribble where the object was. The untouched photo becomes the target. ML Intern wrote and tested pair-building scripts in a CPU sandbox, then ran them as GPU jobs, recording author and license for every source photo. It built 6,042 training pairs and a 160-pair test set, with 40 test pairs from 23 object classes excluded from training entirely.
Before training, it measured the base model alone and with the Viggle turbo LoRA, verifying that batch edits produced identical images to make evaluation cheaper. Training ran 2,000 steps in 1 hour 38 minutes on one A100 (USD 4). Checkpoint comparison on 48 test pairs selected step 500.
Paired with the Viggle turbo LoRA at 6 steps, 67.5% of objects appeared where drawn, at 4.7 seconds per edit. Objects from the 23 unseen classes landed as reliably as the rest (65.0% versus 64.2%). The project took a day and 59 jobs. Total compute: USD 24.
Distilled prompt rewriter
The pocket rewriter at the top of this post started by generating 8,797 short image requests with a small instruct model through Inference Providers. The mix matched the prompt spec: photos, posters, logos, infographics, about a third requesting exact quoted text, many in non-English languages. The 9B teacher rewrote all of them on one A100 in 2 hours 37 minutes (USD 6.50). After quality filtering, 1,840 examples made the training dataset.
Training 0.8B and 2B students took 12 and 18 minutes on an A10G (USD 0.75 combined). The 0.8B ships as an 812 MB GGUF for CPU inference. The project took 11 hours and 24 jobs. Total compute: USD 16.
Four-step image generation
Logolabs' Agate Preview 002 is a 260M-parameter text-to-image model small enough for browsers, but it needs 50 steps with guidance—100 network passes per image. ML Intern distilled it to 4 passes.
The first run cached 155,000 training images as latents, baked guidance into the model, then cut step count in stages from 16 to 8 to 4, all on A100s. The 4-step student beat the teacher run at 4 steps on GenEval and FID. ML Intern exported it to ONNX for browsers. This run took 13 hours and USD 22.
What this changes
ML Intern isn't doing novel ML research. It's orchestrating known techniques—distillation, LoRA training, dataset generation—with full autonomy over compute resources. The barrier dropped from "find a specialist who can dedicate days" to "write a detailed prompt."
The prompts themselves are interesting artifacts. They're executable specifications that combine goal definition, constraint documentation, methodology selection, and cost control. They're version-controlled, shareable, and iteratively refinable. Each prompt encodes not just what to build but how to verify it worked.
The cost structure matters. These projects ran USD 2 to USD 24 in compute. The limiting factor wasn't budget—it was prompt quality. Better problem framing, tighter smoke tests, and clearer success metrics produced better models faster.
The GitHub repository of prompts is effectively a cookbook of ML workflows that execute autonomously. Someone building a similar citrus classifier doesn't start from scratch—they fork a working prompt, swap the dataset, adjust the evaluation metric, and run. The knowledge transfer isn't in trained weights; it's in executable instructions.
This workflow breaks the mental model where custom models require ML specialists. It doesn't eliminate the need for ML expertise—the prompts demonstrate deep understanding of training dynamics, evaluation design, and failure modes. But it shifts that expertise from execution to specification. You still need to know what a good baseline looks like, which metrics matter, and when overfitting starts. You don't need to babysit training runs or debug CUDA errors.
The really interesting question is what happens when these prompts start forking and evolving in public. We've seen this dynamic with code: Stack Overflow answers and GitHub gists become canonical solutions that get copied and refined. If ML project prompts follow the same path, the ML Intern repository becomes infrastructure—a growing library of proven workflows for common modeling tasks.
That's not hype about AI agents. That's observable behavior: an engineer shipped seven production models in a week by writing detailed instructions, and now those instructions are public templates.