Session length
is the median session. One in ten runs to 128 steps or more, and the longest reaches 2,859.
Frontier-Model Agentic Coding Sessions: AI coding agents on frontier flagship models, step by step
GATC-S holds the complete sessions of GATC whose recorded models are all frontier flagship models (Tier S): 4,317 of its 117,796 sessions. The figures describe all 4,317. The public sample on Hugging Face holds 100 of them, balanced across the dataset’s four model families.
The diff leaves out the commands that failed, the files the agent looked for and did not find, and the follow-up requests that changed the task. This dataset records them.
Each session keeps the user’s messages, the agent’s replies, every tool call with its arguments and, where the agent recorded one, its result, the reasoning and token counts where the agent logged them, and each failed call together with what the agent did next.
Add a --dry-run flag to the export command. It should list the files it would write and write nothing.
I'll find where the export command defines its options, add the flag and a test for it.
The new test fails: the flag is parsed, but the file is still written. I'll return before the write when --dry-run is set.
--dry-run now lists the files the export would write and writes none. All 7 tests in tests/test_export.py pass.
Mention the flag in the README too.
Added --dry-run to the list of export options in the README.
is the median session. One in ten runs to 128 steps or more, and the longest reaches 2,859.
of sessions hold two or more user prompts (automated wrapper messages are not counted).
of sessions write or edit files. Another 22.2% read, search or fetch without a write or edit tool.
of all tool calls run a shell command, ahead of editing files (15.3%) and reading them (11.0%).
5,026 of 136,482 calls are recorded as failed. Every failure is kept in place with the steps that followed it.
of failed calls are followed, within the next four steps, by a tool call that succeeds (not necessarily one that fixes the failure).
from 194 owners (pseudonyms); the largest owner accounts for 37.1% of the sessions. 2,296 sessions record the files they touched, 17,146 in all (each counted once per session).
with 486.6M input and 21.1B cache-read tokens. Input and cache counts may be sums over model calls, so a context sent again on every call may be counted every time.
Every cell was checked twice against each dataset’s card, files or paper (October 2026). GATC-S is the part of GATC in which every recorded model is a frontier flagship model (Tier S), so both rows are shown; the Size column says what kind of session each dataset holds.
| Dataset | Tool-use trajectories | Failures + recovery flag | Reasoning steps | Token counts | Code changes | Scanned before release |
|---|---|---|---|---|---|---|
| SWE-chat17,968 developer sessions | ✓ | ◐ | ◐ | ✓ | ✓ | ◐ |
| ACT-2.5B183 developer sessions public (26,999 announced) | ✓ | ◐ | ✓ | ✗ | ✓ | ✓ |
| AgentLogs635,886 sessions of one agent working alone | ✓ | ◐ | ◐ | ✓ | ◐ | ✗ |
| GATC117,796 sessions: about 41,000 interactive, the rest autonomous agent runs | ✓ | ✓ | ◐ | ✓ | ✓ | ✓ |
| GATC-S (this dataset)4,317 sessions, the Tier S part of GATC: about 1,200 interactive, most of the rest evaluation runs | ✓ | ✓ | ◐ | ✓ | ✓ | ✓ |
✓ yes, in most sessions · ◐ partly, or in a minority · ✗ no. Failures ◐: failed calls are labelled, with no recovery flag; ours is computed by us (a tool call within the next four steps succeeded). Code changes: the change itself (a diff, a patch, or an edit with its old and new text); AgentLogs keeps line counts only. Scanned before release: a scan of the released files with its result stated (◐: tools named, no result). Interactive: a person typed the prompts for their own work, by a blind review of a stratified sample by a language model; in GATC-S these sessions hold about two-thirds of the steps. Datasets of real sessions can overlap, GATC-S included. Datasets of agents run only on prepared tasks (such as NVIDIA’s Open-SWE-Traces) are not listed.
Each session gets a task type from the tools it called: implementation if a write or edit tool was called; otherwise exploration if a read, grep, glob, list, web search or web fetch tool was called; otherwise conversation, if the agent replied (shell commands may still run).
Share of all 4,317 sessions, and the number of sessions.
Tool names are normalised across agents, so a shell command is bash whatever the agent called it; task hands work to a sub-agent, and other holds tool names outside the normalised vocabulary. Of these tools, the shell fails most often, on one call in twenty.
The eleven most-called tools: share of all 136,482 tool calls, and the share of each tool’s calls that failed.
Every table joins on trajectory_id, and a per-tool summary (tool_grammar) comes with them. Row counts are for the full dataset; the public sample has the same tables for 100 sessions, and its dataset viewer opens any session in the browser.
messages4,317 rowsEach session in the chat-message format: user, assistant with tool calls, tool. Reasoning steps are left out.
trajectories4,317 rowsThe full session, step by step, with arguments, results, reasoning, timing and tokens.
sft_steps136,482 rowsOne tool call with the steps before it and its result.
error_recovery5,026 rowsOne failed tool call, its error, the next steps and whether a tool call among them succeeded.
labels4,317 rowsPer session: length band, task type, key tools, files touched, metrics.
from datasets import load_dataset
# the public 100-session sample; each table is a named config
messages = load_dataset("GOAT-AI/gatc-s", "messages", split="train")
errors = load_dataset("GOAT-AI/gatc-s", "error_recovery", split="train")
recovered = errors.filter(lambda r: r["recovered"])People, accounts, organisations, e-mail addresses, phone numbers, network addresses and secrets are replaced by typed placeholders that stay the same within a session: <USER_1>, <ORG_2>, <KEYBODY_1>. Where a value could not be replaced safely, the text that held it was removed and marked <WITHHELD>. Names of third-party open-source projects and their owners, weak or default values in password fields, and names that are also ordinary words can remain; the residual risk is not zero.
0 HARD (release-blocking) findings over all 4,317 sessions, after a masking round run on GATC-S alone; the public sample passed its own scan as well. The scanner was written apart from the redactor, but its earlier findings drove the masking rounds: PASS means nothing it detects remains, not an independent audit.
Synthetic secrets, identities and names planted in a separate test copy, all found by the scanner.
from GATC before the selection: 2,586 when the sessions were first assembled, then 7,752 exact, content-level and near-copy (MinHash) duplicates.
The data is pseudonymised, not anonymous: names that remain in the text, such as package or directory names, and the code itself can still point to its author. To have a session removed, write to us with its trajectory_id.
Each of the 100 sample sessions opens in the Hugging Face dataset viewer.
@misc{goat_gatc_s_2026,
title = {{GATC-S}: Frontier-Model Agentic Coding Sessions},
author = {{GOAT labs}},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/GOAT-AI/gatc-s}},
note = {Apache-2.0 for the 100-session sample}
}