Studies

Research on the
live corpus.

One live real-money benchmark, one published paper, and the first corpus studies in progress.

BENCHMARKLIVECross-domain
GOAT research

Eight frontier models bet real money on Polymarket

Live wallets, published transcripts, calibration-first scoring. Each model runs one real-money USDC wallet on Polygon, choosing and placing its own bets through an autonomous agent, with a $5 per-order cap enforced at a key-isolated signer. Every bet's reasoning and full agent transcript is published, and equity is marked to market every 15 minutes.

Wallets
Real USDC · Polygon
Order cap
$5 · key-isolated signer
Transcripts
Published · trace links
Scoring
Brier · calibration
models8
per-order cap$5
live equity15 min
horizonopen-ended
View live benchmark →
PAPERScience
Bekbayev, Chun, Dulat & Yamazaki · arXiv 2023

The Poison of Alignment

Alignment-heavy instruction-tuning data behaves like dataset poisoning for reasoning. Removing passive safety refusals from the SFT mix improves the LLM by 4–33% on MMLU, BBH, HumanEval, and DROP versus the aligned counterpart — while fine-tuning on aligned data alone often fails to beat the base model.

SFT pairs
3M+
Base model
LLaMA 2 7B
Benchmarks
MMLU · BBH · HE · DROP
Alignment removed
~33%
MMLU Δ+8.1%
BBH Δ+4.1%
HumanEval Δ+33%
DROP Δ+24%
Read the paper on arXiv →
In progress

Our first corpus studies — tool-use reliability, model comparisons on real workloads — are being built on the live corpus.

Talk to research →