Pine AI introduced Pine Computer on October 9, 2026: a cloud computer built for AI agents instead of people. Rather than staring at screenshots like a human, the computer signals what changed directly to the model — and on Pine’s benchmark, GPT-5.6 Luna on Pine Computer scored 78.3% on SaaS-Bench v1.1 checkpoints, ahead of Opus 5 with Claude Code (74.3%), at about $1.02 in model cost per task versus $26.50. It’s in private beta for developers.
San Francisco — October 9, 2026
What it is
Pine Computer is virtual computers plus a harness and runtime, reached through an SDK and API. A product creates a Pine Computer when a job needs one, hands it the task, and gets back the files, records, and answers. The core idea: today’s agents work through computers designed for human eyes and hands — they look at a picture of the screen, figure out what’s on it, act, and look again, burning model capability on workarounds.
Pine changed the environment instead. The OS, browser, and applications push change notifications to the model, which Pine describes as “epoll for AI perception”: the model reads change as structure rather than reconstructing it from pixels. Screens still exist for people, who can watch the live view, take control, and hand it back. Each computer runs sealed in its own sandbox, saved state encrypted with the developer’s own key.
The benchmark, honestly
Pine’s paper reports SaaS-Bench v1.1 results (106 tasks across 23 business apps): GPT-5.6 Luna on Pine Computer passed the most checkpoints of nine systems in the published table (78.3%), with a resolved rate of 27.4% — lower than Claude Code’s 31.1%. Checkpoint score measures task progress along the way; resolved rate measures fully finished tasks. So the honest read: more progress per task, at far lower model cost, but fewer end-to-end completions than the leaders.
Why it matters
This is the most interesting agent-infrastructure idea we’ve seen in a while: stop making the model smarter and start making the computer worthy of the model. If structure-instead-of-screenshots generalizes, the cost of running business AI agents could drop an order of magnitude — and cheap agents change what’s worth automating. The reliability gap (27.4% fully resolved) is the thing to watch; Pine says it’s a starting point and plans to publish evaluation tools so others can reproduce and challenge the results. A waitlist is open at pinecomputer.io.
Quick questions
Can I try Pine Computer now? It’s in private beta for developers; you can join the waitlist at pinecomputer.io.
Can I bring my own model? Not yet — that capability is coming, and Pine says a model gets in once it passes its benchmark checklist. It runs Pine’s own model and GPT-5.6 Luna today.
Is it faster and cheaper? On Pine’s SaaS-Bench v1.1 results, GPT-5.6 Luna on Pine Computer made more checkpoint progress at roughly 1/25 the model cost per task of the leading comparisons — but resolved fewer whole tasks. Pine notes these are whole-system comparisons with different budgets, not an isolated test of the computer.
Sources: PR Newswire (Pine AI announcement), Pine Computer launch blog, RuntimeWire

