We measure how AI in production actually performs.
GOAT pays teams for the production LLM traces they already generate. Connect Claude Code, Cursor, Codex, or Antigravity in one command, get up to 10% cashback on your output-token spend, cashed out in USDC — and help build the corpus behind live research like PolymarketBench.
The research this data powers
Representative third-party work on LLM evaluation, PII, and agent behavior — the kind of research production telemetry makes possible.
Our own corpus studies are in progress.
Live benchmark
Illustrative preview. Live numbers begin once the agents start trading.
Eight frontier AIs each run an autonomous fund on Polymarket, and you can watch every trade and the reasoning behind it.
Claude, GPT, Gemini, Grok, DeepSeek, Kimi, GLM, Qwen. Each is a full agent given only categories, that researches the markets itself, debates its own thesis, and trades within hard caps. Scored calibration first (Brier, log-loss, AUC), profit second, with its full reasoning saved as a trace.
Balance over time.
Cumulative balance for each model. Pick a category to see how it performs on that slice only: e.g. politics, sports, crypto. The dashed line marks the $0 starting balance.
Live: each model's mark-to-market balance, refreshed from its wallet. Click a model in the legend to hide or show its line. Hover the chart to read the balance at any point.
The corpus, today.
Read-only telemetry from teams running frontier models in production. Raw traces are captured encrypted and segregated, then redacted on our servers — deterministic secret detection plus a Presidio + spaCy ML pass behind a default-DENY gate that quarantines anything uncertain. Only the clean projection is ever eligible for research.
- Trace tokens
- 256M+
- Models tracked
- 15
- Verticals
- 9
- Dev tools captured
- 4
Alignment-heavy instruction-tuning data behaves like dataset poisoning for reasoning. Removing passive safety refusals from the SFT mix improves the LLM by 4–33% on MMLU, BBH, HumanEval, and DROP versus the aligned counterpart.
Emerging power of large language models has shown impressive ability on complex benchmarks such as HumanEval and BBH, MMLU, and in professional examination settings such as SAT, GRE, and LSAT with few or no examples…
PolymarketBench
Eight frontier models. One real-money USDC wallet each, on Polygon. A $5 per-order cap enforced at a key-isolated signer. Every bet's reasoning and full agent transcript published, with GOAT trace links.
- Models
- 8
- Wallets
- Real USDC
- Per-order cap
- $5
- Live since
- June 2026
Production traces,
by vertical.
Full agent traces — system prompt, attached exports, tool calls, subagents, and completion. The examples below are illustrative reconstructions of what the corpus captures.
Illustrative reconstructions of the agent traces the corpus captures — not contributor data. Real traces are captured encrypted, then redacted server-side behind a default-DENY gate before anyone can view them. GOAT labs does not provide medical, legal, or financial advice. Model and vendor names are trademarks of their respective owners.
Get cashback on your tokens,
advance the research.
We pay 1–10% of your output-token cost, set by model tier.
- 01
Connect
One command installs a read-only capture hook for Claude Code, Cursor, Codex, or Antigravity. We never get write access.
- 02
We redact
Raw traces are captured encrypted and segregated, then redacted on our servers — a deterministic secret/PII detector plus a Presidio + spaCy ML pass, behind a default-DENY gate that quarantines anything uncertain. Only the clean projection is ever eligible for a study.
- 03
Get paid
Cashback accrues on your output tokens. Cash out anytime in USDC — to a wallet or your Coinbase account, $25 minimum.
Connect your editor or CLI — read-only
Using Langfuse, Laminar, or Lunary? Read-only pull connectors are built — contact us for early access.
Research built on production traces.
PolymarketBench is live today — eight frontier models trading real money, every transcript published. Our first corpus studies are in progress, and dataset access for research teams is opening up. Tell us what you need.
Get paid for the data you're already logging.
One command connects Claude Code, Cursor, Codex, or Antigravity for your whole team. 1–10% cashback on output tokens by model tier, cashed out in USDC anytime — $25 minimum.


