Long sessions
runs to 166 steps or more, and the longest reaches 36,656. The median session has 26 steps.
Agentic Coding Sessions: AI coding agents at work, step by step
GATC is the GOAT AI Trajectory Corpus. The figures describe the full dataset of 117,796 sessions. The public sample on Hugging Face holds 1% of it: 1,189 sessions, drawn at random from its 106,012 complete sessions (with a user prompt and an agent reply).
The diff leaves out the commands that failed, the files the agent looked for and did not find, and the follow-up requests that changed the task. This dataset records them.
Each session keeps the user’s messages, the agent’s replies, every tool call with its arguments and, where the agent recorded one, its result, the reasoning and token counts where the agent logged them, and each failed call together with what the agent did next.
Add a --dry-run flag to the export command. It should list the files it would write and write nothing.
I'll find where the export command defines its options, add the flag and a test for it.
The new test fails: the flag is parsed, but the file is still written. I'll return before the write when --dry-run is set.
--dry-run now lists the files the export would write and writes none. All 7 tests in tests/test_export.py pass.
Mention the flag in the README too.
Added --dry-run to the list of export options in the README.
runs to 166 steps or more, and the longest reaches 36,656. The median session has 26 steps.
of all steps and tool calls are in the roughly 41,000 sessions in which a person typed the prompts for their own work (our estimate). Those sessions run long; the others are autonomous agent runs on tasks a script or another agent gave.
of sessions hold two or more user prompts (automated wrapper messages are not counted).
of sessions write or edit files. Another 19.7% read, search or fetch without a write or edit tool.
240,957 of 5,751,020 calls are recorded as failed, each kept in place with the steps that followed it. In 85.5% of them a tool call within the next four steps succeeds (computed by us; not necessarily one that fixes the failure).
with 91,091 calls to MCP servers and 14,318 skill invocations, from sessions of many agents normalised to one schema.
from 3,011 owners (pseudonyms; 10,362 sessions have neither). 60,382 sessions record the files they touched, 550,962 in all (each counted once per session).
with 120.9B input and 543.5B cache-read tokens. Input and cache counts may be sums over model calls, so a context sent again on every call may be counted every time.
Every cell was checked twice against each dataset’s card, files or paper (October 2026). The Size column says what kind of session each dataset holds.
| Dataset | Tool-use trajectories | Failures + recovery flag | Reasoning steps | Token counts | Code changes | Scanned before release |
|---|---|---|---|---|---|---|
| SWE-chat17,968 developer sessions | ✓ | ◐ | ◐ | ✓ | ✓ | ◐ |
| ACT-2.5B183 developer sessions public (26,999 announced) | ✓ | ◐ | ✓ | ✗ | ✓ | ✓ |
| AgentLogs635,886 sessions of one agent working alone | ✓ | ◐ | ◐ | ✓ | ◐ | ✗ |
| GATC (this dataset)117,796 sessions: about 41,000 interactive, the rest autonomous agent runs | ✓ | ✓ | ◐ | ✓ | ✓ | ✓ |
✓ yes, in most sessions · ◐ partly, or in a minority · ✗ no. Failures ◐: failed calls are labelled, with no recovery flag; ours is computed by us (a tool call within the next four steps succeeded). Code changes: the change itself (a diff, a patch, or an edit with its old and new text); AgentLogs keeps line counts only. Scanned before release: a scan of the released files with its result stated (◐: tools named, no result). Interactive: a person typed the prompts for their own work, by a blind review of 360 sessions by a language model; these sessions hold about two-thirds of all steps. Datasets of real sessions can overlap, GATC included. Datasets of agents run only on prepared tasks (such as NVIDIA’s Open-SWE-Traces) are not listed.
Each session gets a task type from the tools it called: implementation if a write or edit tool was called; otherwise exploration if a read, grep, glob, list, web search or web fetch tool was called; otherwise conversation, if the agent replied (shell commands may still run).
Share of all 117,796 sessions, and the number of sessions. The other 2,305 sessions (2.0%) have no agent reply.
Tool names are normalised across agents, so a shell command is bash whatever the agent called it; task hands work to a sub-agent, and other holds tool names outside the normalised vocabulary. Of the common tools, fetching a web page fails most often; the shell fails on one call in eighteen.
The eleven most-called tools: share of all 5,751,020 tool calls, and the share of each tool’s calls that failed.
Each model is shown as a capability tier: S frontier flagship, A frontier, B workhorse, C small and fast, D legacy. A session counts under the highest tier it used.
| Tier | Sessions | Steps per session, median | Tool calls per session, mean | Failed calls | Output tokens per session, median |
|---|---|---|---|---|---|
| S | 4,887 | 23 | 60.1 | 3.3% | 2,976 |
| A | 53,852 | 25 | 65.2 | 3.4% | 2,040 |
| B | 25,450 | 32 | 41.0 | 6.7% | 2,702 |
| C | 13,131 | 25 | 23.5 | 9.0% | 1,429 |
| D | 5,061 | 28 | 20.5 | 1.7% | 6,500 |
102,381 sessions used a model with a tier. Token medians cover the sessions that report output tokens. Tiers differ in agent and task as well as in model, so the columns describe the data, not the models alone.
Every table joins on trajectory_id, and a per-tool summary (tool_grammar) comes with them. Row counts are for the full dataset; the public sample has the same tables at 1% of the size, and its dataset viewer opens any session in the browser.
messages117,796 rowsEach session in the chat-message format: user, assistant with tool calls, tool. Reasoning steps are left out.
trajectories117,796 rowsThe full session, step by step, with arguments, results, reasoning, timing and tokens.
sft_steps5,751,020 rowsOne tool call with the steps before it and its result.
error_recovery240,957 rowsOne failed tool call, its error, the next steps and whether a tool call among them succeeded.
labels117,796 rowsPer session: length band, task type, key tools, files touched, metrics.
from datasets import load_dataset
# the public 1% sample; each table is a named config
messages = load_dataset("GOAT-AI/gatcv6", "messages", split="train")
errors = load_dataset("GOAT-AI/gatcv6", "error_recovery", split="train")
recovered = errors.filter(lambda r: r["recovered"])People, accounts, organisations, e-mail addresses, phone numbers, network addresses and secrets are replaced by typed placeholders that stay the same within a session: <USER_1>, <ORG_2>, <KEYBODY_1>. Where a value could not be replaced safely, the text that held it was removed and marked <WITHHELD>. Names of third-party open-source projects and their owners, weak or default values in password fields, and names that are also ordinary words can remain; the residual risk is not zero.
0 HARD (release-blocking) findings over all 117,796 sessions. The scanner was written apart from the redactor, but its earlier findings drove the masking rounds: PASS means nothing it detects remains, not an independent audit.
Synthetic secrets, identities and names planted in a separate test copy, all found by the scanner.
2,586 when the sessions were first assembled, then 7,752 exact, content-level and near-copy (MinHash) duplicates.
The data is pseudonymised, not anonymous: names that remain in the text, such as package or directory names, and the code itself can still point to its author. To have a session removed, write to us with its trajectory_id.
Each of the 1,189 sample sessions opens in the Hugging Face dataset viewer.
@misc{goat_agentic_coding_sessions_2026,
title = {{GATC}: Agentic Coding Sessions},
author = {{GOAT labs}},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/GOAT-AI/gatcv6}},
note = {Apache-2.0 for the 1\% sample}
}