Analysis
Hark previewed Handoff, its browser-use agent, in a demonstration reported by [TechCrunch](https://techcrunch.com/2026/08/05/hark-previews-its-browser-use-agent-for-completing-tasks/). The pitch is narrow and concrete: point it at a website that has no API and it completes the task anyway, reading page structure and pixels to decide where to click and what to type. Demos included ordering flowers, booking a restaurant through OpenTable, shopping on Target and Walmart, and pulling research off LinkedIn. A waitlist is open with a release targeted for later this summer.
The technical claim is the interesting part. Rather than wrapping a general chat model in a scaffolding loop, Hark says it post-trained a model whose output space is actions -- a click at a location, a keyboard input into a field -- instead of text tokens. Full pre-training on that objective is planned for later in 2026. If that holds up, it is a different bet from the scaffolding approach most competitors run, and it should show up as lower latency and lower cost per completed task rather than higher benchmark scores.
Brett Adcock is not a first-time founder raising on a demo. He founded Figure, the humanoid robotics company, and Archer Aviation before that. Hark's $700M Series A in May 2026 was extraordinary for a company at this stage, and it bought the thing that matters for this problem: enough compute and enough human demonstration data to post-train an action model rather than prompt around a general one.
“Full pre-training on that objective is planned for later in 2026.”
The competitive field is crowded and getting more so. TechCrunch names Browser Use, Polar, Strawberry and Aside among startups, and all three frontier labs are shipping in the category -- OpenAI with Operator-lineage products, Anthropic with computer use, and Google with agentic browsing in Gemini. The strategic question for every one of them is the same: are browser agents a feature of the model, or a product? Labs are betting feature. Hark is betting product, with a specialized model underneath.
The bear case is distribution, not capability. Target, Walmart, LinkedIn and OpenTable did not consent to being automated, and the standard countermeasures -- bot detection, rate limits, terms-of-service enforcement, CAPTCHA escalation -- are cheap for them and expensive for Hark. LinkedIn in particular has litigated aggressively against automated access for a decade. A demo that works today can be blocked by a Cloudflare rule tomorrow, and Hark has no contractual relationship to fall back on.
There is also the unglamorous reliability problem. Browser agents fail on the long tail: a modal that appears one time in fifty, a checkout flow that changes mid-quarter, a payment step that needs a real card. Consumers will forgive a chatbot that is wrong; they will not forgive an agent that buys the wrong thing. Whatever Hark's benchmark numbers are, the number that decides the company is the completion rate on unseen sites, unassisted, at the end of a multi-step task.
What to watch: whether Hark publishes a reproducible benchmark against GPT-5.5 and Opus 4.8 rather than a claim, whether any of the named retailers strikes a partnership instead of a block, and what per-task pricing looks like at launch. A browser agent that costs more per completed checkout than the merchant's affiliate commission has no business model, no matter how good the model is.