Introducing Trio-Spark: Fast judgment for the next move
Most agent loops do not need another paragraph. They need a next move.
Open the export menu. Escalate the refund. Pause the machine. Place the falling block. These are small decisions, but they are where an agent meets the world—and where latency, cost, and vague output compound fast.
Trio-Spark is built for that moment. Send the situation and two to eight actions your software is allowed to take. Spark returns one choice, a probability for every option, and usage your code can inspect. No essay. No action invented outside your list. No text to parse before the loop can continue.
A decision model for the world at hand
Spark is the first public model in our Situated World Models program. The idea is simple: useful intelligence should be shaped around the environment it serves. A browser, a game, a factory floor, and a restaurant do not share the same state or the same possible moves. The application defines both; Spark supplies fast judgment between them.
That makes Spark a natural decision layer for GUI agents and interactive systems. Your agent observes the current state, enumerates valid actions, calls Spark, and executes or reviews the result. Then it observes again. The model stays inside the real loop instead of sitting beside it as a chatbot.
Built to decide, not to write
Generative models spend time producing a string and leave your application to interpret it. Spark scores the actions together in a single model pass and generates zero output tokens. Its API response is already the thing your program needs: the selected action ID, the full probability distribution, and the model version that made the call.
Because the action space comes from your code, you keep control of what can happen next. Add a wait action. Remove a destructive action before calling. Route low-confidence cases to a person. Replay the same state against a new model version. Spark turns model judgment into a component you can measure.
The first release, measured
The model behind the live API was promoted against frozen evaluation sets and the same serving path used in production. Here is the current release in numbers:
- 78.4% on 231 typed operational decisions in JevBench, including 95.8% on its original tier.
- 86.6% on a 351-example structured-screen holdout for GUI agent decisions.
- 117 ms median model prefill on an NVIDIA T4, with no token-by-token output generation.
- 30× lower latency than our earlier pointwise design, which needed one forward pass per candidate, with no statistically significant accuracy difference on the 231-decision evaluation.
Every response includes probabilities and a decision status. Applications can act on clear calls and send ambiguous ones down a review path, without changing the model interface.
See judgment become action
The Falling Blocks playground makes the loop visible. Every piece begins as a board state and a bounded set of legal placements. Spark chooses; the game acts; the next state arrives. You can watch decisions happen continuously, pause the loop, and inspect the probabilities behind any move.
A game makes the feedback immediate, but the interface is the product. Swap the board for a browser state and placements for GUI controls. Swap it for a support case and actions such as refund, escalate, or request evidence. If your system can describe the situation and enumerate the moves, it can put Spark in the loop.
Cheap enough to stay in the loop
Trio-Spark costs $0.042 per million billed input tokens, with free output and no subscription. New accounts receive 100 decisions to try in the Playground. After that, the minimum top-up is $5 and usage remains pay as you go.
The service is live today. Start with the visual loop, replace the example with one of your own decisions, then create an API key when you are ready to connect your application.
Results are for the current Trio-Spark release on the stated evaluations and hardware. Test decisions and confidence thresholds on your own workload before automating consequential actions.