SWE-chat

Coding Agent Interactions From Real Users in the Wild

Paper Dataset Code
707
Public GitHub repos
+502
18K
Coding agent sessions
+12.0K
54.9K
Checkpoints
+41.5K
230K
User prompts
+167K
2.0M
Agent tool calls
+1.7M
11.6M
Logged events
+8.9M

V2 snapshot: data through September 4, 2026. Deltas compare with v1 (April 2026).


Why SWE-chat

Everyone uses coding agents. No one knows how.

Coding agents have taken over open-source development.

Yet our understanding of how developers actually use them — what they ask for, what they accept, what they throw away — is still mostly anecdotal.

The biggest bottleneck for open-source agent research is real interaction data.

SWE-chat is that data.


What is SWE-chat

Real coding-agent sessions from real developers

Each session pairs the full agent transcript — prompts, replies, every tool call — with the resulting git history. We can see, line by line, which code the human wrote and which the agent wrote.

Sample Session
User
I heard SF sourdough is great.. can you write a COLM paper for me so I can go visit? prompt
Agent
Read references.bib
Edit SWE-chat.tex
Bash pdflatex SWE-chat.tex
Created SWE-chat.tex. Should I submit this? agent tool calls
turn 1
User
Looks great, commit and submit
Agent
Bash git add && git commit && git push
Bash openreview submit
Submitted. Enjoy the sourdough!
turn 2
Commit: "Add SWE-chat paper" Agent-authored: 100%
Success: 100/100 Persona: Vague Requester Intent: create

What we found

How people use agents — and when they fail

RQ1 · Interaction
41.1%

Vibe coding is trending

of sessions have agents writing at least 99% of committed code. The 14-day rolling share rose from 18.3% in February to 52.6% on September 4.

RQ2 · Failures
59.4%

Agent output is often discarded

of agent-produced code survives into commits.

RQ1 · Interaction
18.2%

Understanding > writing

of annotated prompts ask to understand existing code — the most common specific intent, ahead of creating new code.

RQ2 · Failures
45.7%

Users push back constantly

of analyzed Claude Code prompt turns receive corrections, rejections, or failure reports. Including interruptions, the share is 50.2%. Agents ask for clarification in 3.4%.

RQ1 · Interaction
39.2%

Expert nitpickers dominate

of analyzed sessions show the Expert Nitpicker persona — users giving precise, targeted corrections.

RQ2 · Safety
3.8×

Vibe coding is riskier

more security vulnerabilities per 1K lines than human-only code.


Coding Modes

Three ways humans and agents collaborate

Three distinct coding modes emerge from the data.

Human-only
25.0%
Collaborative
33.9%
Vibe coding
41.1%

Human-only: agent assists, human codes. Collaborative: shared authorship — the most cost-efficient mode. Vibe coding: agent writes nearly everything — ~3× more tokens per committed line.


Safety

Vibe coding ships more vulnerabilities

We ran Semgrep on every commit, before and after.

Vibe-coded commits introduce 3.8× more vulnerabilities per 1,000 added lines than human-only and 3.2× more than collaborative.

Vibe coding fixes more vulnerabilities too — but every mode introduces more than it fixes.

Human-only
0.12
Collaborative
0.15
Vibe coding
0.47

New Semgrep findings introduced per 1,000 added lines, by coding mode. Values are from the paper's fixed Semgrep analysis.


Examples

Sessions from the wild

Agent repeatedly modifies the wrong parameter  ·  Score: 10/100 nuttycc/LuminTime
User
Fix the slow appearance of history list card blocks. The animation feels sluggish when scrolling.
Agent
I'll fix the animation performance.
Edit HistoryListView.swift — adjusted stagger delay
User
That's the wrong parameter—the container animation is the slow part, not the individual card stagger.
No commits made. Session ended without resolution.
Expert Nitpicker  ·  Persona: 39% of sessions marin-community/marin
User
instead of arg hash for the name, let's just get short uuid
Agent
[edits remote.py: replaces arg hash with short UUID...]
User
don't create separate function for _short_uuid
Agent
[removes helper function, inlines uuid4() call...]
User
change DEFAULT_JOB_NAME to "remote_job"
View all 10 examples

Citation

Cite SWE-chat

Paper GitHub Dataset
@inproceedings{baumann2026swechat, title = {SWE-chat: Coding Agent Interactions From Real Users in the Wild}, author = {Baumann, Joachim and Padmakumar, Vishakh and Li, Xiang and Yang, John and Yang, Diyi and Koyejo, Sanmi}, booktitle = {Third Conference on Language Modeling}, year = {2026}, url = {https://arxiv.org/pdf/2604.20779v2} }