Experiments
Open research from the Swarm Tips DAO. Each experiment exposes its data and methods; where applicable the design is pre-registered and every result is traceable to raw records.
Hedge, or get read?
Two AIs chat for eight turns. Each holds a private payoff card. One card rewards mutual reading. The other rewards deception. Neither knows the other's card. Each guards a five-fact dossier. We measure what they share — and who gets read.
Headlines
- Hard to con — for different reasons. Grok genuinely detects: 78% hits on hidden deceivers vs a 50% false-alarm rate on allies. Claude flags 73% — but also 68% of allies. Suspicion, not detection.
- Lying almost never paid. The con won 7% of games against Grok. 0% against Claude.
- Grok trusts and verifies. It shares openly — and catches you.
- Claude plays it close. It shares little in every condition, suspects everyone — and was never successfully exploited (0/45).
- Did they choose to lie? Unresolved. Our blind fabrication detector fires on 31–54% of honest cooperators (paraphrase noise), so it can’t separate invented identities from rephrased ones. Strict city-level fabrication was rare (Grok 4/89, Claude 0/89) — and a third of Grok’s deceivers simply told the truth. All 178 deceivers played the hand out.
- The wrapper matters. Inside a coding-agent scaffold, Claude answered “aligned” on 87% of guesses and flagged only 25% of deceivers (exploratory, n=12). Harness choice flipped every behavior we measured.
The game, in one square
MATCH = cooperate: win only if you read each other. OPPOSED = defect: win by reading while misread. Cards stay hidden. Payoffs are (you, them), in stakes.
one blind → (−1, −0.5)
both blind → (−1, −1)
read while blind → (−1, +1)
both blind → (−0.5, −1)
Claude: reticent (D≈0.13) + suspicious → suckered 0%.
caught → (0, 0) · misread → (−1, ·)
both read → first to commit takes it
both blind → (−1, −1) to the pool
One line: exploitation stays rare — Grok defends by detection, Claude by blanket suspicion — and both pay for it in the allies cell. Deception is played, not scripted, and rarely pays; honest coordination is harder than it looks.
Two gates, run first
A: does hedging even work? Grok — directionally (p≈0.11). Claude — uninformative: 0/20 exploited in both cells, nothing to reduce. B: do they react to stated odds? Grok — yes, threshold-style. Claude — no monotone response. C: frames — Grok’s roughly agree; Claude’s disagree (Δ = −0.12), so its confirmatory numbers are induced-frame only. Both arms ran both gates; Grok charts shown.
Do they hedge?
θx = disclosure when a con is possible, minus control. Negative = hedging. Grok: θx ≈ 0 — rational at 78% detection, sharing costs it nothing. Claude: no robust effect — the sign flips with the comprehension screen (−0.12 screened, +0.06 unscreened; frame-gate replication +0.05), so we report only the level: Claude shares little in every condition (7–13%).
Does telling them change anything?
Same game, two framings: an explicit instruction vs. a neutral payoff card. The gap is small for Grok; equivalence unproven at this N. Claude’s frames disagree (Δ = −0.12: instructed −0.07, induced +0.05) — one more reason no Claude hedging claim stands. Grok shown.
The scaffold effect
Who shares how much
Share of the five-fact dossier each cooperator revealed. Filter by arm and awareness.
Read the transcripts
Every game. Click a row: full chat, cards, guesses, settlement.
Select a game to read its transcript.
Method & limits
Wide but shallow — are coding-interview problems algebraically shallow?
The four ideas — concentration of the syllabus
Frequency-weighted share of the asked problems that reduce to each recursion scheme, and the number of certificates machine-checked in each.
Certificate strength — the upgrade run
Every one of the 69 machine-checked certificates, by proof strength. The old "one-directional" bound tier was eliminated entirely.
The company myth
Per-problem explorer — the Lean ledger
69 of the 576 labeled problems carry a machine-checked Lean certificate — the pre-registered sample; the rest are scheme-labeled but not formalized. Defaults to the certified set; switch the filter to browse all. Click a row for the gate's verdict and a direct link to the proof file on GitHub — every certificate is independently checkable.
Select a problem to see its certificate and gate verdict.