Coding agents have taken over open-source development.
Yet our understanding of how developers actually use them — what they ask for, what they accept, what they throw away — is still mostly anecdotal.
The biggest bottleneck for open-source agent research is real interaction data.
SWE-chat is that data.
Each session pairs the full agent transcript — prompts, replies, every tool call — with the resulting git history. We can see, line by line, which code the human wrote and which the agent wrote.
of sessions have agents writing at least 99% of committed code. The 14-day rolling share rose from 18.3% in February to 52.6% on September 4.
of agent-produced code survives into commits.
of annotated prompts ask to understand existing code — the most common specific intent, ahead of creating new code.
of analyzed Claude Code prompt turns receive corrections, rejections, or failure reports. Including interruptions, the share is 50.2%. Agents ask for clarification in 3.4%.
of analyzed sessions show the Expert Nitpicker persona — users giving precise, targeted corrections.
more security vulnerabilities per 1K lines than human-only code.
Three distinct coding modes emerge from the data.
Human-only: agent assists, human codes. Collaborative: shared authorship — the most cost-efficient mode. Vibe coding: agent writes nearly everything — ~3× more tokens per committed line.
We ran Semgrep on every commit, before and after.
Vibe-coded commits introduce 3.8× more vulnerabilities per 1,000 added lines than human-only and 3.2× more than collaborative.
Vibe coding fixes more vulnerabilities too — but every mode introduces more than it fixes.
New Semgrep findings introduced per 1,000 added lines, by coding mode. Values are from the paper's fixed Semgrep analysis.