GATC

Agentic Coding Sessions: AI coding agents at work, step by step

Coding-agent sessions
117,796
Steps
9.7M
Tool calls
5.75M
Interactive sessions
~41K
Failed tool calls
241K
Reasoning steps
572K

GATC is the GOAT AI Trajectory Corpus. The figures describe the full dataset of 117,796 sessions. The public sample on Hugging Face holds 1% of it: 1,189 sessions, drawn at random from its 106,012 complete sessions (with a user prompt and an agent reply).

Why this dataset

A final diff shows the result. A session shows how the agent got there.

The diff leaves out the commands that failed, the files the agent looked for and did not find, and the follow-up requests that changed the task. This dataset records them.

What is in a session

Every step, in the order it happened

Each session keeps the user’s messages, the agent’s replies, every tool call with its arguments and, where the agent recorded one, its result, the reasoning and token counts where the agent logged them, and each failed call together with what the agent did next.

Illustration of the format, not a session from the dataset
User

Add a --dry-run flag to the export command. It should list the files it would write and write nothing.

Agent

I'll find where the export command defines its options, add the flag and a test for it.

  • grep add_argument cli/
  • read /home/<USER_1>/src/exporter/cli/export.py
  • edit /home/<USER_1>/src/exporter/cli/export.py
  • edit /home/<USER_1>/src/exporter/tests/test_export.py
  • bash python -m pytest tests/test_export.py -q (failed)1 failed, 6 passed. test_dry_run_writes_nothing: AssertionError: out/report.csv was written …
Agent

The new test fails: the flag is parsed, but the file is still written. I'll return before the write when --dry-run is set.

  • edit /home/<USER_1>/src/exporter/cli/export.py
  • bash python -m pytest tests/test_export.py -q

--dry-run now lists the files the export would write and writes none. All 7 tests in tests/test_export.py pass.

User

Mention the flag in the README too.

Agent
  • edit /home/<USER_1>/src/exporter/README.md

Added --dry-run to the list of export options in the README.

14 steps8 tool calls1 failed, recoveredimplementation
What the data shows

How agents work, and where they fail

Sessions
1 in 10

Long sessions

runs to 166 steps or more, and the longest reaches 36,656. The median session has 26 steps.

Interaction
68%

Steps in interactive sessions

of all steps and tool calls are in the roughly 41,000 sessions in which a person typed the prompts for their own work (our estimate). Those sessions run long; the others are autonomous agent runs on tasks a script or another agent gave.

Turns
38.1%

Multi-turn sessions

of sessions hold two or more user prompts (automated wrapper messages are not counted).

Tasks
45.4%

Nearly half of sessions edit files

of sessions write or edit files. Another 19.7% read, search or fetch without a write or edit tool.

Failures
1 in 24

Failed calls are kept

240,957 of 5,751,020 calls are recorded as failed, each kept in place with the steps that followed it. In 85.5% of them a tool call within the next four steps succeeds (computed by us; not necessarily one that fixes the failure).

Agents
99,088

Hand-offs to sub-agents

with 91,091 calls to MCP servers and 14,318 skill invocations, from sessions of many agents normalised to one schema.

Breadth
3,939

Projects

from 3,011 owners (pseudonyms; 10,362 sessions have neither). 60,382 sessions record the files they touched, 550,962 in all (each counted once per session).

Tokens
1.8B

Output tokens, as the agents reported them

with 120.9B input and 543.5B cache-read tokens. Input and cache counts may be sums over model calls, so a context sent again on every call may be counted every time.

Comparison

How it differs from other agent datasets

Every cell was checked twice against each dataset’s card, files or paper (October 2026). The Size column says what kind of session each dataset holds.

DatasetTool-use trajectoriesFailures + recovery flagReasoning stepsToken countsCode changesScanned before release
SWE-chat17,968 developer sessions✓◐◐✓✓◐
ACT-2.5B183 developer sessions public (26,999 announced)✓◐✓✗✓✓
AgentLogs635,886 sessions of one agent working alone✓◐◐✓◐✗
GATC (this dataset)117,796 sessions: about 41,000 interactive, the rest autonomous agent runs✓✓◐✓✓✓

✓ yes, in most sessions · ◐ partly, or in a minority · ✗ no. Failures ◐: failed calls are labelled, with no recovery flag; ours is computed by us (a tool call within the next four steps succeeded). Code changes: the change itself (a diff, a patch, or an edit with its old and new text); AgentLogs keeps line counts only. Scanned before release: a scan of the released files with its result stated (◐: tools named, no result). Interactive: a person typed the prompts for their own work, by a blind review of 360 sessions by a language model; these sessions hold about two-thirds of all steps. Datasets of real sessions can overlap, GATC included. Datasets of agents run only on prepared tasks (such as NVIDIA’s Open-SWE-Traces) are not listed.

Task types

What the sessions do

Each session gets a task type from the tools it called: implementation if a write or edit tool was called; otherwise exploration if a read, grep, glob, list, web search or web fetch tool was called; otherwise conversation, if the agent replied (shell commands may still run).

Implementationa write or edit tool was called
45.4%
53,516
Conversationno write, edit, read or search tool; shell and other tools may run
32.9%
38,808
Explorationread, search or fetch; no write or edit
19.7%
23,167

Share of all 117,796 sessions, and the number of sessions. The other 2,305 sessions (2.0%) have no agent reply.

Tools

Which tools agents call, and how often each one fails

Tool names are normalised across agents, so a shell command is bash whatever the agent called it; task hands work to a sub-agent, and other holds tool names outside the normalised vocabulary. Of the common tools, fetching a web page fails most often; the shell fails on one call in eighteen.

ToolShare of callsFailed
bash53.4%5.6%
read14.3%2.6%
edit11.3%3.1%
other5.3%2.3%
grep3.4%1.0%
todo2.9%0.2%
write2.6%2.7%
task1.7%1.4%
mcp1.6%6.0%
glob1.2%2.4%
web_fetch0.5%10.3%

The eleven most-called tools: share of all 5,751,020 tool calls, and the share of each tool’s calls that failed.

Model tiers

Sessions by the model tier they used

Each model is shown as a capability tier: S frontier flagship, A frontier, B workhorse, C small and fast, D legacy. A session counts under the highest tier it used.

TierSessionsSteps per session, medianTool calls per session, meanFailed callsOutput tokens per session, median
S4,8872360.13.3%2,976
A53,8522565.23.4%2,040
B25,4503241.06.7%2,702
C13,1312523.59.0%1,429
D5,0612820.51.7%6,500

102,381 sessions used a model with a tier. Token medians cover the sessions that report output tokens. Tiers differ in agent and task as well as in model, so the columns describe the data, not the models alone.

The data

Five tables, joined on one key

Every table joins on trajectory_id, and a per-tool summary (tool_grammar) comes with them. Row counts are for the full dataset; the public sample has the same tables at 1% of the size, and its dataset viewer opens any session in the browser.

messages117,796 rows

Each session in the chat-message format: user, assistant with tool calls, tool. Reasoning steps are left out.

trajectories117,796 rows

The full session, step by step, with arguments, results, reasoning, timing and tokens.

sft_steps5,751,020 rows

One tool call with the steps before it and its result.

error_recovery240,957 rows

One failed tool call, its error, the next steps and whether a tool call among them succeeded.

labels117,796 rows

Per session: length band, task type, key tools, files touched, metrics.

Load the 1% sample
from datasets import load_dataset

# the public 1% sample; each table is a named config
messages = load_dataset("GOAT-AI/gatcv6", "messages", split="train")
errors = load_dataset("GOAT-AI/gatcv6", "error_recovery", split="train")
recovered = errors.filter(lambda r: r["recovered"])
De-identification

Checked before release

People, accounts, organisations, e-mail addresses, phone numbers, network addresses and secrets are replaced by typed placeholders that stay the same within a session: <USER_1>, <ORG_2>, <KEYBODY_1>. Where a value could not be replaced safely, the text that held it was removed and marked <WITHHELD>. Names of third-party open-source projects and their owners, weak or default values in password fields, and names that are also ordinary words can remain; the residual risk is not zero.

PASS
Release gate verdict

0 HARD (release-blocking) findings over all 117,796 sessions. The scanner was written apart from the redactor, but its earlier findings drove the masking rounds: PASS means nothing it detects remains, not an independent audit.

1,657 / 1,657
Scanner self-test

Synthetic secrets, identities and names planted in a separate test copy, all found by the scanner.

10,338
Duplicates removed

2,586 when the sessions were first assembled, then 7,752 exact, content-level and near-copy (MinHash) duplicates.

The data is pseudonymised, not anonymous: names that remain in the text, such as package or directory names, and the code itself can still point to its author. To have a session removed, write to us with its trajectory_id.

Examples

Sample sessions on Hugging Face

Each of the 1,189 sample sessions opens in the Hugging Face dataset viewer.

Citation

Cite the dataset

BibTeX
@misc{goat_agentic_coding_sessions_2026,
  title        = {{GATC}: Agentic Coding Sessions},
  author       = {{GOAT labs}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/GOAT-AI/gatcv6}},
  note         = {Apache-2.0 for the 1\% sample}
}