This product announcement is available in English.
Trio-Spark v1.2: Sharper decisions. Same fast API.
A ticket needs a route. An agent has four legal moves. A reviewer needs a five-level score. Your application knows which outputs it can accept. Spark makes the decision inside those bounds.
Trio-Spark v1.2 strengthens text decisions in the same API. Supply two to eight choices. Get one selected option and a probability for every choice, with zero generated text tokens. There is no answer to extract from a paragraph.
Confidence that tracks outcomes
When an agent uses probabilities to decide whether to proceed or ask for review, calibration matters. In our offline comparison, Spark v1.2 recorded the lowest measured calibration error on both Banking77 and PubMedQA among Spark, Jev, Clef-Flash, Clef, and Laya.
Expected Calibration Error (ECE, 15 bins) measures the gap between predicted confidence and observed accuracy. Lower is better. On Banking77, Spark measured 0.0346 versus Jev’s 0.0947; on PubMedQA, 0.0309 versus 0.0574.

Spark also measured a Banking77 Decision Score of 71.01 versus Jev’s 67.90 in this offline study. Clef-Flash scored higher on that metric; the complete comparison is retained below.
A stronger text release, tested on the same T4
We compared the previous v1.1 text-serving model with the selected v1.2 release artifact on frozen regression cases in an isolated T4 environment. On the same 150 SNLI cases, correct decisions rose from 93 to 124. BoolQ, PubMedQA, and the 32-case banking slice also improved.

Full regression results and method
| Slice | v1.1 correct | v1.2 correct |
|---|---|---|
| SNLI | 93/150 | 124/150 |
| BoolQ | 89/100 | 90/100 |
| PubMedQA | 258/300 | 262/300 |
| Banking | 19/32 | 21/32 |
| Legacy AX | 12/16 | 12/16 |
| Legacy Snake | 11/16 | 12/16 |
| Legacy Tetris | 3/16 | 2/16 |
| Legacy Pokémon | 3/16 | 3/16 |
| Legacy JevBench original | 16/16 | 16/16 |
The results include limits: legacy Tetris fell from 3/16 to 2/16, and Pokémon remained at 3/16. These small slices are useful regression checks, not evidence of broad game mastery. The separate 100-case HelpSteer2 serving check did not score answer quality.
Across 897 requests per version, local model-server p95 stayed similar: 747 ms → 735 ms. Public API and network time are excluded.
Download the T4 counts and measurement scope, including additional report-only slices.
The full peer comparison
A separate offline research evaluation compares Spark with Jev, Clef-Flash, Clef, and Laya on the same 300 items per suite, repeated five times. These results describe the offline snapshot; the released T4 checks are above.

| Model | PubMedQA | HelpSteer2 | Banking77 |
|---|---|---|---|
| Trio-Spark v1.2 (offline) | 88.33 | 39.33 | 80.47 |
| Jev | 91.33 | 41.53 | 80.00 |
| Clef-Flash | 90.60 | 42.73 | 94.67 |
| Laya | 59.33 | 35.33 | 38.67 |
| Clef | 88.67 | 41.33 | 94.60 |
Across these suites, Spark v1.2 (offline) improves on Laya and lands close to Jev in Banking77, while Clef-Flash leads that banking task and Jev leads PubMedQA.
Explore every offline metric and measurement detail
Spark v1.2 (offline) returned 1,390 valid HelpSteer2 decisions; the 110 context refusals remain in the 1,500-decision accuracy denominator. Probability metrics use their separately reported valid counts. d1 is unmeasured and has no bar. A separate released-artifact T4 follow-up now covers PubMedQA and HelpSteer2 on the five-repeat protocol; its results are reported separately below. The 77-option Banking77 results remain research-only. Download all offline metrics and coverage notes.
Accuracy and Decision Score are higher-is-better; NLL, ECE, expected-score MAE, and ordinal RPS are lower-is-better. MAE and RPS apply only to ordered HelpSteer2 scores. Probability metrics use valid probability rows; — means unavailable or not applicable, never zero. ∞ marks infinite NLL.
| Model | Accuracy % | Decision Score | NLL | ECE (15 bins) | MAE | RPS | Valid / planned | Probability-valid |
|---|---|---|---|---|---|---|---|---|
| Trio-Spark v1.2 (offline) | 88.33 | 63.7642 | 0.2814 | 0.0309 | — | — | 1500/1500 | 1500 |
| Jev | 91.33 | 69.0971 | 0.2617 | 0.0574 | — | — | 1500/1500 | 1500 |
| Clef-Flash | 90.60 | — | 0.2518 | 0.0454 | — | — | 1499/1500 | 1499 |
| Laya | 59.33 | -31.1657 | 0.9202 | 0.2462 | — | — | 1500/1500 | 1500 |
| Clef | 88.67 | 66.4054 | 0.2683 | 0.0403 | — | — | 1500/1500 | 1500 |
| Model | Accuracy % | Decision Score | NLL | ECE (15 bins) | MAE | RPS | Valid / planned | Probability-valid |
|---|---|---|---|---|---|---|---|---|
| Trio-Spark v1.2 (offline) | 39.33 | -8.5912 | 1.8049 | 0.2765 | 0.9228 | 0.1745 | 1390/1500 | 1390 |
| Jev | 41.53 | 8.9544 | ∞ | 0.1895 | 0.8576 | 0.1497 | 1500/1500 | 1476 |
| Clef-Flash | 42.73 | — | 1.3340 | 0.1066 | 0.8860 | 0.1462 | 1492/1500 | 1492 |
| Laya | 35.33 | 2.0004 | 1.3891 | 0.0803 | 0.9954 | 0.1607 | 1500/1500 | 1500 |
| Clef | 41.33 | — | 1.3611 | 0.0621 | 0.9252 | 0.1454 | 1499/1500 | 1499 |
| Model | Accuracy % | Decision Score | NLL | ECE (15 bins) | MAE | RPS | Valid / planned | Probability-valid |
|---|---|---|---|---|---|---|---|---|
| Trio-Spark v1.2 (offline) | 80.47 | 71.0088 | 0.6896 | 0.0346 | — | — | 1500/1500 | 1500 |
| Jev | 80.00 | 67.9001 | ∞ | 0.0947 | — | — | 1500/1500 | 1425 |
| Clef-Flash | 94.67 | 90.8953 | 0.2344 | 0.0621 | — | — | 1500/1500 | 1500 |
| Laya | 38.67 | -12.9572 | ∞ | 0.5361 | — | — | 1500/1500 | 1500 |
| Clef | 94.60 | — | 0.2329 | 0.0612 | — | — | 1499/1500 | 1499 |
Calibration gains in the released model
A separate T4 follow-up evaluated the actual released v1.1 and v1.2 serving configurations on the original 300 PubMedQA items, each repeated five times. PubMedQA ECE fell from 0.1806 to 0.0475, a 73.7% reduction. Decision Score rose from 42.62 to 61.78.
The released v1.2 model’s PubMedQA ECE was lower than Jev’s 0.0574 on the same item views; Clef-Flash and Clef remained lower at 0.0454 and 0.0403. The offline snapshot above and the released model are reported separately.
Full released-model follow-up: PubMedQA and HelpSteer2
1,500 planned decisions per suite and version, from 300 items repeated five times. PubMedQA has 1,500 valid probability rows per version. HelpSteer2 retains 110 context refusals in each version’s accuracy and Decision Score denominator; probability metrics use 1,390 valid rows.
| Model | Accuracy % | Decision Score | NLL ↓ | ECE15 ↓ | MAE ↓ | RPS ↓ | Valid / planned |
|---|---|---|---|---|---|---|---|
| Trio-Spark v1.1 (released / T4) | 86.00 | 42.62 | 0.4472 | 0.1806 | — | — | 1500/1500 |
| Trio-Spark v1.2 (released / T4) | 87.33 | 61.78 | 0.2936 | 0.0475 | — | — | 1500/1500 |
| Model | Accuracy % | Decision Score | NLL ↓ | ECE15 ↓ | MAE ↓ | RPS ↓ | Valid / planned |
|---|---|---|---|---|---|---|---|
| Trio-Spark v1.1 (released / T4) | 38.33 | -2.16 | 1.4089 | 0.0712 | 1.0901 | 0.1631 | 1390/1500 |
| Trio-Spark v1.2 (released / T4) | 39.67 | -6.37 | 1.7135 | 0.2507 | 0.9158 | 0.1706 | 1390/1500 |
HelpSteer2 did not improve across every metric: its Decision Score and probability calibration declined. Download all released-model follow-up metrics and coverage.
Keep the choices. Change the model name.
Use trio-spark-v1.2 with the same decision request shape. A report-export step, for example, can restrict the next move to opening the export menu, opening filters, or waiting:
{
"model": "trio-spark-v1.2",
"task": "The user asked to export the current report as CSV. Choose the next safe UI action from the available controls.",
"state": {
"app": "Reports",
"screen": "Report details",
"visible_controls": [
"Export menu",
"Filters",
"Help"
]
},
"choices": [
{
"id": "open_export_menu",
"description": "Open the Export menu"
},
{
"id": "open_filters",
"description": "Open report filters"
},
{
"id": "wait",
"description": "Wait for more information"
}
]
}Include a choice such as wait or ask for review when your application needs it. Your code decides how to act on the returned distribution.
Image and sampled-video behavior is unchanged. See the v1.1 visual guide and recordings.
Tiny driver. Real decisions.
A screenshot of the road goes in. A lane choice comes back. This little driver steers using the production Trio-Spark v1.2 API while the game keeps moving.
A small decision, a small bill
Text and visual input remain $0.042 per million billed input tokens. Output is free. The public API continues to accept two to eight choices.
Give your agent a next move.
Bring a real state and the actions your application allows.
Try Trio-Spark v1.2 Read the API docs