Terminal agents · Agent harness · Test-time scaling

Mid-Harness

Scaling Actions Between Model and Harness for Terminal Agents

Minki Kang1,2* Ryo Hachiuma1 Shaokun Zhang1 Subhashree Radhakrishnan1 Yonggan Fu1 Jindong Jiang1 Mingjie Liu1 Ehsan Hosseini-Asl1 Yi Dong1 Yu-Chiang Frank Wang1 Byung-Kwan Lee1†

1NVIDIA 2KAIST

*Work done during internship †Project lead

Headline results

Same model. Same harness. Every score goes up.

Pass@1 and Pass@3 (%) of the base agent and Mid-Harness in four transfer settings
Benchmark Agentmodel × harness · verifier Pass@1base → Mid-Harness Pass@3base → Mid-Harness
TerminalBench-Lite
Nemotron3.5 Lightning30BTerminus-2zero-shot
+2.04pp41.16 → 43.20 +3.06pp55.10 → 58.16
Terminal-Bench 2.1
Nemotron3 Ultra550BTerminus-2zero-shot
+5.24pp50.94 → 56.18 +1.12pp65.17 → 66.29
TerminalBench-Lite
Qwen3.5-9BTerminus-2zero-shot
+2.04pp40.48 → 42.52 +1.02pp60.20 → 61.22
SWE-bench-VerifiedMini · 50 tasks
TMAX-9BVanillux2distilled
+2.00pp46.67 → 48.67 +8.00pp54.00 → 62.00

Pass@1 / Pass@3 (%) over three runs per task; large numbers are gains over the base agent in percentage points. Mid-Harness uses pairwise verification with N = 8: zero-shot on the Terminus-2 rows (no verifier training; the Nemotron verifiers are the same backbone), distilled on SWE-bench-Verified. All seven transfer settings

Illustrative · multi-turn · two agents
User task

Make the failing tests pass.

base agent1 sample

t2in detail

  1. 01
    Sampleπ(·|ht) × 1
    $pip install yaml
  2. 02
    Verifyskipped

    The first sample goes straight to the harness.

  3. 03
    Executeharness → environment
    $pip install yaml→ runs in the env
✗ Error compounds · yaml still missing
mid-harnessN = 8 → verify → 1

t2in detail

  1. 01
    Sampleπ(·|ht) × 8
  2. 02
    Verifyψ 8 ring + 16 pivot duels, before execution
  3. 03
    Execute oneothers discarded
    $pip install pyyaml→ runs in the env
✓ Verified every turn · 12 tests pass

Findings at a glance

Sample N actions. Verify them. Run only the best one.

Generator coverage

The generator already proposes better actions.

Base agent50.00
+ frontier verifier68.03

GPT-5.6 Sol picks among TMAX-9B's own eight candidates — the generator is not trained.

Verification matters

How candidates are compared decides the gain.

Listwise51.02
Pairwise54.76
Distilled pairwise57.14

Self-verification with TMAX-9B; distillation from GPT-5.6 Sol trains only the verifier.

Complements trajectory scaling

Action scaling stacks with trajectory scaling.

Best-of-T55.10
+ Mid-Harness66.33
Sequential Refine55.10
+ Mid-Harness60.20

Same number of environment runs: Best-of-T with T = 3, SR with R = 1.

TMAX-9B Pass@1 (%) on TerminalBench-Lite · Mid-Harness uses N = 8 · bars start at 45%

Verification without reasoning

Verification may not need reasoning.

A verifier that answers each duel with a single token, A or B, scored higher than one that reasons first, and cost less.

One duel · two verifiers Apip install -e . vs Bpip install pyyaml

Schematic: a verifier that writes out its reasoning and a verifier that answers with one token compare the same two actions and pick the same one.

1tokenper duel, no rationale
59.18%Pass@1, up from 57.14% with reasoning
−24.1%end-to-end cost

Distilled TMAX-9B verifier on TerminalBench-Lite, N = 8; at N = 4, the verifier that reasons still scores higher. Duel is schematic.

01 · How it works

Verify, then act.

Mid-Harness samples candidates, verifies them, and passes one action to the unchanged harness. The generator is not trained.

Step 01 · Sample

Draw N candidates

The model proposes several next actions from the same history.

Step 02 · Verify

Compare before running

A verifier checks the candidates and picks one.

Step 03 · Execute one

Only the winner runs

The harness executes it; the others are discarded.

Three ways to verify

Same N candidates, three ways for the verifier ψ to pick one.

PairwiseRing duels pick 4 pivots, who then face the rest — not all 28 pairs. 22–25 callsPass@1 54.76%

Pass@1 with the TMAX-9B generator and a zero-shot TMAX-9B verifier on TerminalBench-Lite, N = 8. Diagrams are schematic.

02 · When does action scaling work?

Results

With the TMAX-9B generator and its Vanillux2 harness fixed, we vary candidate width, verification mechanism, and verifier capability on TerminalBench-Lite (98 tasks, three runs each).

2.1

A strong verifier reveals coverage; width alone does not

Candidate width · TMAX-9B

Verification mechanism determines the return from more candidates

N = 1 is the base agent without action scaling. The frontier verifier (GPT-5.6 Sol) uses listwise verification because of its cost. × marks a first-runnable-candidate proxy at N = 8.

03 · Cost and transfer

Scaling

Best-of-T and Sequential Refine need fresh environment runs. Mid-Harness improves both without adding any, and its gains carry over to other models, benchmarks, and harnesses.

3.1

Action scaling vs. parallel scaling

Token cost · verifier without reasoning (one token) · TMAX-9B · TerminalBench-Lite

More Pass@1 per dollar than parallel scaling

5.8× cheaperMid-Harness (Distilled, N = 8) matches Best-of-7 at 59.18% Pass@1: $0.16 vs. $0.91 per run.

24% cheaperMid-Harness (Distilled, N = 8) costs $0.16 instead of $0.21 with reasoning, and Pass@1 rises from 57.14% to 59.18%.

+6.13 ptsMid-Harness (Distilled) + Best-of-3 beats Best-of-7, 65.31% vs. 59.18%, at 33% lower cost.

The verifier answers each duel with one token, A or B. At N = 4 this is cheaper but scores lower (52.72% and 54.42%). From the paper's appendix.

Token cost · verifier with reasoning · TMAX-9B · TerminalBench-Lite

Also cheaper with a verifier that reasons first

3× cheaperMid-Harness (Distilled, N = 8) matches Best-of-5 at 57.14% Pass@1: $0.21 vs. $0.62 per run.

4× cheaperMid-Harness (Distilled, N = 4) edges out Best-of-3, 55.44% vs. 55.10%: $0.08 vs. $0.34.

+7.15 ptsMid-Harness (Distilled) + Best-of-3 beats Best-of-7, 66.33% vs. 59.18%, at 12% lower cost.

Mid-Harness uses pairwise verification that writes out its reasoning before each verdict.

Estimated token cost per run at OpenRouter Qwen3.5-9B rates ($0.08 / $0.13 per million input / output tokens), not a measured serving bill. Points are operating points, not an equal-budget experiment.
3.2

Transfer across models, benchmarks, and harnesses

Model scale · TerminalBench-Lite

Action scaling helps from 4B to 27B

Pass@1 / Pass@3 over three runs per task; both Mid-Harness variants use pairwise verification with N = 8. “Best Pass@1 gain” is the larger of the evaluated Mid-Harness variants' gains over the base agent.

04 · What still limits verification?

Analysis

Offline, on 21 held-out tasks, we measure how often pairwise verifiers agree with the GPT-5.6 Sol teacher. This is agreement, not action correctness.

Verification agreement by turn

Gains persist, but later states stay harder

1,355 states with valid comparisons for both models (83.4% of states). The 33+ bin contains only eight tasks.

Clear verifier failures

Command semantics and feasibility dominate what remains

Teacher disagreements that GPT-5.6 Terra judges clear verifier failures, by category: 3,328 for zero-shot and 1,810 after distillation.

Open problem

Judge a command without running it.

Even after distillation, most clear verifier failures come from misjudging what a command will do in the current environment.

67.4%

of the distilled verifier's clear failures

  • 39.0% candidate semantics
  • 28.4% execution feasibility

Candidate semantics 39.0%

What does this code actually do?

token[4] = 'a' + s1;

The verifier flagged this as a bug and said it should be s1 + 'a', but both compute the same letter.

Execution feasibility 28.4%

Can it actually run here?

if '9090' in cmdline or 'main.go' in cmdline:
    os.kill(int(pid), 9)

The verifier rewarded the port cleanup and persistent launch, missing that it can kill itself first.

Distilled TMAX-9B verifier vs. the GPT-5.6 Sol reference, offline; from the paper's appendix.