Trio-Spark v1.2 · October 8, 2026

Trio-Spark v1.2: Sharper decisions. Same fast API.

A ticket needs a route. An agent has four legal moves. A reviewer needs a five-level score. Your application knows which outputs it can accept. Spark makes the decision inside those bounds.

Trio-Spark v1.2 strengthens text decisions in the same API. Supply two to eight choices. Get one selected option and a probability for every choice, with zero generated text tokens. There is no answer to extract from a paragraph.

Confidence that tracks outcomes

When an agent uses probabilities to decide whether to proceed or ask for review, calibration matters. In our offline comparison, Spark v1.2 recorded the lowest measured calibration error on both Banking77 and PubMedQA among Spark, Jev, Clef-Flash, Clef, and Laya.

Expected Calibration Error (ECE, 15 bins) measures the gap between predicted confidence and observed accuracy. Lower is better. On Banking77, Spark measured 0.0346 versus Jev’s 0.0947; on PubMedQA, 0.0309 versus 0.0574.

Offline ECE15, lower is better. Banking77: Spark 0.0346, Jev 0.0947, Clef-Flash 0.0621, Clef 0.0612, Laya 0.5361. PubMedQA: Spark 0.0309, Jev 0.0574, Clef-Flash 0.0454, Clef 0.0403, Laya 0.2462.
Offline research snapshot; 300 items per suite, each repeated five times. ECE uses valid probability rows, whose counts differ by model; full counts and other metrics are below. Banking77 uses 77 research options, beyond the public API’s two-to-eight-choice limit. These are measured slice results, not an overall accuracy ranking.

Spark also measured a Banking77 Decision Score of 71.01 versus Jev’s 67.90 in this offline study. Clef-Flash scored higher on that metric; the complete comparison is retained below.

A stronger text release, tested on the same T4

We compared the previous v1.1 text-serving model with the selected v1.2 release artifact on frozen regression cases in an isolated T4 environment. On the same 150 SNLI cases, correct decisions rose from 93 to 124. BoolQ, PubMedQA, and the 32-case banking slice also improved.

Matched T4 regression: SNLI 93 to 124 of 150; BoolQ 89 to 90 of 100; PubMedQA 258 to 262 of 300; Banking 19 to 21 of 32.
Selected regression slices, measured October 8, 2026. Counts share the same cases within each row; they are not an overall benchmark score.
Full regression results and method
Full primary and legacy T4 regression counts
Slicev1.1 correctv1.2 correct
SNLI93/150124/150
BoolQ89/10090/100
PubMedQA258/300262/300
Banking19/3221/32
Legacy AX12/1612/16
Legacy Snake11/1612/16
Legacy Tetris3/162/16
Legacy Pokémon3/163/16
Legacy JevBench original16/1616/16

The results include limits: legacy Tetris fell from 3/16 to 2/16, and Pokémon remained at 3/16. These small slices are useful regression checks, not evidence of broad game mastery. The separate 100-case HelpSteer2 serving check did not score answer quality.

Across 897 requests per version, local model-server p95 stayed similar: 747 ms → 735 ms. Public API and network time are excluded.

Download the T4 counts and measurement scope, including additional report-only slices.

The full peer comparison

A separate offline research evaluation compares Spark with Jev, Clef-Flash, Clef, and Laya on the same 300 items per suite, repeated five times. These results describe the offline snapshot; the released T4 checks are above.

Offline research accuracy comparison across PubMedQA, HelpSteer2, and research-only 77-option Banking77. Spark v1.2 (offline) scores 88.33%, 39.33%, and 80.47%, respectively.
Banking77 uses 77 research options; the public API supports two to eight choices. Peer systems use their respective API or SDK paths. This is not a matched hardware or latency comparison.
Offline strict accuracy (% of all 1,500 planned decisions per suite)
ModelPubMedQAHelpSteer2Banking77
Trio-Spark v1.2 (offline)88.3339.3380.47
Jev91.3341.5380.00
Clef-Flash90.6042.7394.67
Laya59.3335.3338.67
Clef88.6741.3394.60

Across these suites, Spark v1.2 (offline) improves on Laya and lands close to Jev in Banking77, while Clef-Flash leads that banking task and Jev leads PubMedQA.

Explore every offline metric and measurement detail

Spark v1.2 (offline) returned 1,390 valid HelpSteer2 decisions; the 110 context refusals remain in the 1,500-decision accuracy denominator. Probability metrics use their separately reported valid counts. d1 is unmeasured and has no bar. A separate released-artifact T4 follow-up now covers PubMedQA and HelpSteer2 on the five-repeat protocol; its results are reported separately below. The 77-option Banking77 results remain research-only. Download all offline metrics and coverage notes.

Accuracy and Decision Score are higher-is-better; NLL, ECE, expected-score MAE, and ordinal RPS are lower-is-better. MAE and RPS apply only to ordered HelpSteer2 scores. Probability metrics use valid probability rows; — means unavailable or not applicable, never zero. ∞ marks infinite NLL.

PubMedQA
ModelAccuracy %Decision ScoreNLLECE (15 bins)MAERPSValid / plannedProbability-valid
Trio-Spark v1.2 (offline)88.3363.76420.28140.0309——1500/15001500
Jev91.3369.09710.26170.0574——1500/15001500
Clef-Flash90.60—0.25180.0454——1499/15001499
Laya59.33-31.16570.92020.2462——1500/15001500
Clef88.6766.40540.26830.0403——1500/15001500
HelpSteer2
ModelAccuracy %Decision ScoreNLLECE (15 bins)MAERPSValid / plannedProbability-valid
Trio-Spark v1.2 (offline)39.33-8.59121.80490.27650.92280.17451390/15001390
Jev41.538.9544∞0.18950.85760.14971500/15001476
Clef-Flash42.73—1.33400.10660.88600.14621492/15001492
Laya35.332.00041.38910.08030.99540.16071500/15001500
Clef41.33—1.36110.06210.92520.14541499/15001499
Banking77 · research only
ModelAccuracy %Decision ScoreNLLECE (15 bins)MAERPSValid / plannedProbability-valid
Trio-Spark v1.2 (offline)80.4771.00880.68960.0346——1500/15001500
Jev80.0067.9001∞0.0947——1500/15001425
Clef-Flash94.6790.89530.23440.0621——1500/15001500
Laya38.67-12.9572∞0.5361——1500/15001500
Clef94.60—0.23290.0612——1499/15001499

Calibration gains in the released model

A separate T4 follow-up evaluated the actual released v1.1 and v1.2 serving configurations on the original 300 PubMedQA items, each repeated five times. PubMedQA ECE fell from 0.1806 to 0.0475, a 73.7% reduction. Decision Score rose from 42.62 to 61.78.

The released v1.2 model’s PubMedQA ECE was lower than Jev’s 0.0574 on the same item views; Clef-Flash and Clef remained lower at 0.0454 and 0.0403. The offline snapshot above and the released model are reported separately.

Full released-model follow-up: PubMedQA and HelpSteer2

1,500 planned decisions per suite and version, from 300 items repeated five times. PubMedQA has 1,500 valid probability rows per version. HelpSteer2 retains 110 context refusals in each version’s accuracy and Decision Score denominator; probability metrics use 1,390 valid rows.

PubMedQA · released artifacts, same T4
ModelAccuracy %Decision ScoreNLL ↓ECE15 ↓MAE ↓RPS ↓Valid / planned
Trio-Spark v1.1 (released / T4)86.0042.620.44720.1806——1500/1500
Trio-Spark v1.2 (released / T4)87.3361.780.29360.0475——1500/1500
HelpSteer2 · released artifacts, same T4
ModelAccuracy %Decision ScoreNLL ↓ECE15 ↓MAE ↓RPS ↓Valid / planned
Trio-Spark v1.1 (released / T4)38.33-2.161.40890.07121.09010.16311390/1500
Trio-Spark v1.2 (released / T4)39.67-6.371.71350.25070.91580.17061390/1500

HelpSteer2 did not improve across every metric: its Decision Score and probability calibration declined. Download all released-model follow-up metrics and coverage.

Keep the choices. Change the model name.

Use trio-spark-v1.2 with the same decision request shape. A report-export step, for example, can restrict the next move to opening the export menu, opening filters, or waiting:

{
  "model": "trio-spark-v1.2",
  "task": "The user asked to export the current report as CSV. Choose the next safe UI action from the available controls.",
  "state": {
    "app": "Reports",
    "screen": "Report details",
    "visible_controls": [
      "Export menu",
      "Filters",
      "Help"
    ]
  },
  "choices": [
    {
      "id": "open_export_menu",
      "description": "Open the Export menu"
    },
    {
      "id": "open_filters",
      "description": "Open report filters"
    },
    {
      "id": "wait",
      "description": "Wait for more information"
    }
  ]
}

Include a choice such as wait or ask for review when your application needs it. Your code decides how to act on the returned distribution.

Image and sampled-video behavior is unchanged. See the v1.1 visual guide and recordings.

Tiny driver. Real decisions.

A screenshot of the road goes in. A lane choice comes back. This little driver steers using the production Trio-Spark v1.2 API while the game keeps moving.

35 seconds of continuous gameplay. Ten applied model decisions, two coins, one bonk. Screenshots are the model input; no hidden road map or scripted avoidance. Code · Recording details · Download GIF.

A small decision, a small bill

Text and visual input remain $0.042 per million billed input tokens. Output is free. The public API continues to accept two to eight choices.

Give your agent a next move.

Bring a real state and the actions your application allows.

Try Trio-Spark v1.2 Read the API docs