The Interface Problem No One's Talking About
Most agentic models are trapped. GUI-focused models go blind without a screen. Tool-calling models freeze when the app has no API. But real work doesn't respect these boundaries—a single business task often demands clicking through a UI, writing some code, and hitting an API endpoint.
H company just shipped Holo4, a new series of agentic models that refuses to pick sides. The 27B dense and 35B-A3B Mixture of Experts variants interact with software through whatever interface fits: GUIs, code sandboxes, MCP servers, or REST APIs. Same model, same API call, different execution environment.
This matters because the current generation of computer-use agents are specialists masquerading as generalists. You swap models depending on whether you're automating desktop apps or orchestrating cloud services. Holo4 collapses that distinction.
Benchmark Performance vs. Frontier Models
On OSWorld 2.0—the hardest academic benchmark for desktop control—Holo4-27B scores 61.7% while Holo4-35B-A3B reaches 30.9%. Claude Opus 5.5 leads at 81.8%, but Holo4 delivers competitive performance with orders of magnitude fewer parameters.
The cost story is more interesting. H company open-sourced every trajectory behind their benchmark scores, letting anyone replay each agent step at trajectories.hcompany.ai or download them from Hugging Face. They price Holo4-27B on the H Models API based on actual token usage, and the per-task economics are compelling compared to frontier models.
On AutomationBench (API workflow orchestration), Holo4 also competes with much larger closed models at substantially lower cost per completed task. The exact numbers vary by harness version and task subset, but the pattern holds: competitive capability at a fraction of the inference expense.
Real Software, Real Tasks
The benchmark scores matter less than what these models can actually do. H company trained Holo4 on environments and tasks from their "Agentic Task Factory"—a pipeline that generates about 10,000 interactive tasks from documentation alone, including hybrid environments that expose the same state through both a GUI and MCP.
They show Holo4-27B building a detailed 3D model of the Eiffel Tower in FreeCAD from an exacting written spec. The prompt specifies half-widths at different heights, smooth curves that fall steeply near ground and gently higher up, four tapered legs, three platform slabs, a mast. The model produces a solid with correct dimensions, using 84 calls and 1.3M tokens. The base model Qwen3.8 27B attempts the same task but produces a visibly different result in 60 calls.
Another example: building a Pac-Man game in Godot that runs unattended with a heuristic-driven player, scoring pellets and avoiding ghosts. Holo4-27B delivers working code in 68 calls (2.4M tokens, 268 lines). The base model takes 197 calls, burns 11.4M tokens, and writes 327 lines—though it's unclear from the screenshots whether it actually works.
These aren't contrived demos. They're messy, multi-step workflows on professional software where success requires remembering context across hundreds of actions, debugging failures, and adapting strategy mid-task.
How They Built It
Agentic Task Factory
H company's task generation pipeline takes documentation—screenshots of real websites, open-source software guides—and produces interactive environments with verifiable completion criteria. It's generated around 10,000 tasks so far across web apps, MCP servers, desktop environments, and hybrid setups.
This approach sidesteps the usual data bottleneck for agentic training. Instead of manually scripting thousands of workflows, they automate environment and task creation from existing docs. The resulting distribution covers professional software people actually use.
Training Harness Rebuild
Alongside training, they rebuilt the execution harness—the loop that runs the model's actions and manages context over hundreds of steps. Engineers reviewed why agents failed on OSWorld 2.0, and the agents themselves tagged failure reasons.
The two biggest changes: giving the agent reliable memory to track hundreds of steps, and adding a shell on the desktop machine itself. These aren't model improvements—they're infrastructure improvements that let the model's capabilities actually manifest in long workflows.
Holotron4 Nano
As part of the NVIDIA Nemotron Coalition, H company applied their post-training stack to Nemotron 3 Nano Omni, producing Holotron4 Nano. Same recipe, smaller foundation model.
The adapted model shows significant absolute percentage-point improvements over the base model on GUI workflows and environments exposing MCP, APIs, or code sandboxes. This demonstrates the post-training recipe transfers across model sizes and isn't specific to the 27B–35B range.
The Interface Generalization Bet
Most agent builders are specializing. Anthropic focuses on computer use via screenshots and actions. OpenAI emphasizes function calling and structured outputs. Smaller labs pick a niche—browser automation, code execution, RAG pipelines.
H company is betting the opposite direction: a single model that routes to whatever interface fits the subtask. Need to fill a web form? Click and type. Need to query a database? Write SQL. Need to check inventory via API? POST to the endpoint.
This creates different failure modes. A specialist model fails gracefully within its domain. A generalist might choose the wrong interface—trying to GUI-automate something that has a perfectly good API, or calling a nonexistent tool instead of just clicking through the UI.
But if it works, you collapse three agent stacks into one. You stop maintaining separate GUI-agent, code-agent, and tool-agent systems. You stop routing tasks to the right specialist upfront. The model figures out the routing.
Open Weights, Open Trajectories
Both Holo4-27B and Holo4-35B-A3B are available on the H Models API today. Weights are on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF formats, alongside Holotron4 Nano.
Every trajectory behind their public benchmark scores is open-sourced. You can replay each step or download the dataset. This level of transparency is rare—most labs share final scores but not the full execution traces.
H company says optimized DSpark drafter checkpoints are coming soon to further accelerate inference. Speculative decoding with a smaller draft model can meaningfully reduce latency and cost on long agentic runs, especially when the agent is generating verbose reasoning traces or exploring multiple action candidates.
What This Means for Agentic AI
Holo4 represents a different architectural bet than most of the field. Instead of pushing a single interface paradigm to its limits, it trains for interface fluency.
The open question is whether real-world performance matches the benchmark results. Academic benchmarks test capabilities in isolation. Production systems fail due to compounding errors, context confusion, and unexpected edge cases.
But the cost-performance ratio matters. If Holo4-27B can handle 70% of the tasks that Opus 5.5 solves, at 10% of the inference cost, a lot of automation suddenly becomes economically viable. Not every workflow needs frontier model quality.
The Agentic Task Factory also matters strategically. If H company can generate high-quality training environments faster than competitors can manually curate them, they can iterate on a broader task distribution and adapt to new software faster.
We'll see if interface generalization beats specialization. But the early results are compelling enough that other labs will likely start training multi-interface agents. The era of picking one modality is ending.