One of the stupidest things we keep doing with computer-use agents is pretending the screen is the computer.
Need to update a record? Click through six tabs.
Need to download some data? Watch the model squint at a tiny export button.
Need to transform a file? Apparently the correct interface is now a mouse pointer being driven around like a drunk beetle.
Sometimes the GUI is genuinely the only interface available. Fine. Use it.
But if there is an API, a tool call or a bit of code that can do the job directly, making the agent poke pixels is technological cosplay.
That is why H Company's Holo4 release is interesting.
One agent, several ways to touch the world
Holo4 is a family of computer-use models designed to move across interfaces rather than being married to one. H Company says the same model can click and type in graphical interfaces, write and execute code, and call MCP or API tools across desktop, web, Android and sandbox environments.
That is the right architectural instinct.
Real work does not arrive nicely sorted into "GUI tasks" and "API tasks". A single workflow might start in a browser, discover a document, use an API to fetch structured records, run code to clean them and then return to a GUI because some cursed enterprise application only exposes the final approval button through a web page designed in 2014.
A useful agent should choose the best available interface at each stage.
Humans already do this. If I can copy a file with one command, I do not open Explorer, navigate through seven folders and drag the thing while whispering encouragement to Windows. The machine should be allowed the same dignity.
The release is unusually inspectable
H Company has released two main Holo4 variants: a 27B dense model and a 35B-A3B mixture-of-experts model. The Holo4-27B model card lists a 262,144-token maximum context and describes GUI, code and tool-call interfaces.
More importantly, the release includes weights in several formats and public trajectories behind its benchmark runs.
That last bit matters.
Computer-use benchmark numbers are particularly easy to turn into marketing wallpaper because an aggregate score hides all the weirdness: which task version was used, what harness drove the model, whether retries were allowed, how screenshots were presented, what tooling was exposed and whether a failure was the model's fault or the environment having a small nervous breakdown.
H Company itself notes that task releases, subsets and harnesses vary.
If you can inspect and replay trajectories, you can ask better questions than "number big?". You can see whether the model used a sensible strategy, recovered from errors, chose a tool well or simply got lucky.
Open weights do not mean free-for-all
There is an important bit of fine print here. The Holo4-27B model card currently lists a CC-BY-NC-4.0 licence.
So yes, the weights are available. No, that does not mean somebody can automatically stuff them into any commercial product and carry on like licensing is a decorative footer.
There is also the usual hardware reality. Twenty-seven billion parameters is not a rounding error. Quantised formats help, and the mixture-of-experts variant activates fewer parameters at a time, but "downloadable" and "cheap to run at scale" remain different statements.
Open models are wonderful. GPUs still send invoices.
Benchmarks are the beginning, not the verdict
H Company reports strong results for Holo4 and says the models improve substantially over their base models. It also reports that Holo4 trails stronger closed systems on some long-workflow benchmarks.
That is a healthier claim than pretending one leaderboard row has solved computer use.
Long-running agent workflows fail in boring places. Login expires. A modal appears. A file takes longer than expected. The API schema changes. The browser tab gets redirected. The agent performs step 19 correctly and then forgets why step 20 exists.
Those failures are why I care more about interface flexibility and observable trajectories than a victory lap over a single benchmark.
If an agent can recognise that clicking through a table is stupid and switch to an API, that is useful. If it can use code for the transformation and then return to the GUI only for the thing that actually requires the GUI, that is a better system design.
Computer use should not mean screen use
The phrase "computer-use agent" has been too closely associated with models looking at screenshots and pretending to be a person with a mouse.
That was a necessary phase because GUIs are the universal fallback. They let agents operate software that was never designed for agents at all.
But the screen should be the fallback, not the religion.
The interesting direction is an agent that can ask: what is the safest, fastest and most reliable interface for this step?
Sometimes that will be a click.
Sometimes it will be MCP.
Sometimes it will be an API call.
Sometimes it will be five lines of code that save three minutes of watching a cursor wander around a dashboard like it has forgotten where it parked.
Holo4 does not prove that problem is solved.
It does, however, point the agent in the right bloody direction.