AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can Holo4 Support Generalist AI Agents On Computers? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

H Company has released Holo4, a pair of open-weight models designed to handle computer tasks through graphical interfaces, code, MCP and APIs. The company reports a 61.7% OSWorld 2.0 score for its 27B model, but the results have not been independently verified and comparisons use different evaluation setups.

H Company has released Holo4, a series of open-weight models designed to operate software through screens, code and tool interfaces, with one model handling different methods within a task, as detailed in the original analysis. The company reports that its 27B dense model scored 61.7% on OSWorld 2.0; that result is self-reported and has not yet been independently verified.

The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company says both are available through its H Models API and for download from Hugging Face in FP16, FP8 and GGUF formats. The models are intended to click and type in graphical interfaces, write and run code, and call MCP or API tools, selecting an interface based on the task.

H Company reports OSWorld 2.0 scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. It compares the first score with 81.8% for Opus 5.5, which the company identifies as the strongest closed model in its comparison. The release material says the Holo4 scores use the company’s own harness and notes that model releases, harnesses and task subsets differ across comparisons.

The company says it trained Holo4 with supervised and reinforcement learning across a range of environments and tasks, including tasks generated by its Agentic Task Factory. It has also published the trajectories behind its public benchmark results for review. H Company says examples in FreeCAD and Godot show gains over its Qwen base models, though those demonstrations do not establish how reliably Holo4 performs on varied business workflows.

At a glance
announcementWhen: Announced September 2026; independent e…
The developmentH Company released two open-weight Holo4 models for computer-use tasks and published benchmark trajectories alongside its performance claims.
At a glance
announcementWhen: announced via Hugging Face and company…
The developmentH Company announced the release of Holo4, a two-model series of open-weight computer-use agents, along with an updated Holotron4 Nano and open-sourced benchmark trajectories.

One Model Across Software Interfaces

Many computer tasks combine actions across a screen, code and connected tools. A model built for only one interface may be unable to continue when a task moves from clicking through an application to using an API. H Company’s multi-interface design targets that boundary by letting a single model use whichever access method fits each step.

If the reported results hold up in independent evaluations, open weights could give developers more choice over where and how they run software agents, including the option to self-host. H Company also presents Holo4 as less expensive than leading closed models, but that comparison depends on pricing and evaluation assumptions described by the company. The benchmark score alone does not establish the reliability or total cost of running the models on business tasks.

Publishing trajectories gives researchers and developers material to inspect alongside the headline scores. That may help them assess how the model reached an answer and reproduce parts of the evaluation. It does not, by itself, confirm the scores or show that performance will transfer to other software and workflows.

From Holo1 to Holo4

H Company positions Holo4 as a continuation of its work on agentic models, following Holo1, and announced it alongside Holotron4 Nano, an update to Holotron 3. The benchmark notes identify the underlying models as Qwen3.8 27B for Holo4 27B and Qwen3.6 35B-A3B for the MoE version.

The company says many agent models specialize in either graphical interfaces or tool calls. Holo4 is intended to work across desktops, the web, Android, a code sandbox and business APIs, with the same model invocation across platforms. These are stated design goals; the announcement does not provide independent evidence of performance across every listed environment.

H Company’s cost comparisons use its own API rates for Holo4 and Alibaba Cloud list prices for Qwen, with a cache-price assumption for the MoE model. Comparisons involving GPT and Opus use effort sweeps from OpenAI launch data. The company cautions that different releases, harnesses and task subsets limit direct comparisons.

“Real work is not siloed that way, and a single business task can require combining these different approaches.”

— H Company

Benchmark Results Await Review

The published benchmark figures are H Company’s own results, not independent evaluations. The company says its OSWorld 2.0 setup differs from those used for other models, and that the compared releases and task subsets are not uniform. It has published trajectories, but outside replication has not yet established whether the reported scores hold under other evaluation conditions.

The large gap between the two Holo4 models on OSWorld 2.0 — 61.7% for the 27B model and 30.9% for the 35B-A3B model — is not explained in the announcement. Holo4 also has not been evaluated on AutomationBench’s private set. For that benchmark, the company says public-set scores and private-set cost figures are drawn from different evaluation sources. It has not provided third-party evidence of reliability on routine business tasks.

Private Tests and Outside Replication

H Company says it will report Holo4’s results on the AutomationBench private set after that evaluation is complete. Independent submissions and reproductions can also test the published OSWorld 2.0 results against other setups. Developers can access the models through the H Models API or download them from Hugging Face; how they perform in broader use remains to be seen.

Key Questions

What is Holo4?

Holo4 is H Company’s series of open-weight models for computer-use tasks. The company says the models can work through graphical interfaces, code, MCP and APIs.

Which Holo4 models are available?

The release includes a 27B dense model and a 35B-A3B Mixture of Experts model. H Company lists them on its API and on Hugging Face in FP16, FP8 and GGUF formats.

How did Holo4 score on OSWorld 2.0?

H Company reports scores of 61.7% for Holo4 27B and 30.9% for Holo4 35B-A3B. The results use the company’s harness and have not been independently verified.

Have independent evaluators confirmed the results?

Not in the supplied source material. H Company has published trajectories behind its public benchmark scores, which others can inspect, but independent reproductions are still pending.

What remains unknown about Holo4?

The reason for the score gap between the two models is not explained. Their reliability on varied business workflows, and how results compare under consistent independent testing, also remain unclear.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Xiaomi MiMo V2.6

Xiaomi has released MiMo v2.6, a new version of its wireless communication technology, prompting increased industry attention amid limited confirmed details.

24 Jev Ideas For Modeling Decisions In AI

A Sept. 29 article maps 24 uses for Jev, a tool for small AI decisions, and says 15 are live or strong fits under a four-condition test.

A Spooky Guide To AI-Assisted Halloween Design

A practical guide to AI-assisted Halloween design: separating real AI features from motion-sensor props, planning effects, and managing privacy and reliability.

10 Best Mini PCs For Local AI Projects To Explore In 2026

A 2026 mini PC roundup compares memory, storage and expansion options for local AI, but the supplied source identifies only seven models.