Trio-Spark v1.1 · Visual decisions

Trio-Spark v1.1: See the scene. Choose the next move.

A hand signal changes. A person enters the frame. A situation crosses from one allowed action to another. The next move is often visible before it can be written down.

Trio-Spark v1.1 can see that moment and choose what comes next. Bring text, one image, or a short sampled window to one model, one API key, and one wallet. Define the moves your application allows. Spark returns one choice and a probability for every option, without writing an answer first.

Bring the world at hand

Text, images, and sampled camera windows now meet in the same decision product. An existing loop can add what is visible while keeping the same bounded choice contract: the situation and allowed moves go in; one move and the full distribution come back.

In the Playground, start with an uploaded image, a local video, or your camera. The gesture demo makes the loop tangible: show the scene, define the gestures that matter, and watch Spark choose between them.

Choices stay in your application

Spark chooses only from the two to eight actions you supply. Your application can include options such as wait, ask for review, or a small set of domain actions, then decide how to use the returned probabilities. The result is ready for code to inspect rather than prose for code to interpret.

Try a visual decision

The new gesture demo shows the loop end to end. Choose the visual input mode, allow camera access or upload media, and define the gesture choices. The result includes the selected choice, the full distribution, and usage.

Open the gesture demo or read the API docs for image and sampled-video request examples.

Recorded decisions from our deployed visual model. IPN Hand footage, licensed CC BY 4.0. Code and recording details.

Follow a scene in motion

A short sampled window can show how a scene changes. Here, Spark follows movement through Shibuya Crossing and chooses which road users are visibly moving through the center.

Recorded decisions from our deployed visual model. Footage by Basile Morin, licensed CC BY-SA 4.0. Code and recording details.

Simple usage-based pricing

Trio-Spark v1.1 costs $0.042 per million billed input tokens for text and visual requests. The Playground includes 100 free decisions, with no card required. API responses report input-token usage so teams can measure each workflow directly.

Developer details

Use trio-spark-v1.1 for text, one JPEG or PNG image, or two to four timestamped frames from a window of up to 30 seconds. The API docs cover the request schema, media bounds, context limits, errors, and idempotent retries. Our published text benchmarks and recorded game demos remain clearly labeled v1.0 results.

Build the next move

Start with a narrow decision your application can verify, define the allowed actions, and bring the state as text, one image, or a short sampled window.

Try Trio-Spark v1.1