Format guide

Where Claude Code stores sessions — and how to read the JSONL

The format is undocumented and changes without notice. This is what’s actually in the files, including the two traps that produce wrong token numbers.

Updated August 2026 · verified against current Claude Code and Codex CLI output

The paths

# Claude Code — one file per session
~/.claude/projects/<encoded-path>/<session-uuid>.jsonl
~/.claude/projects/<encoded-path>/<session-uuid>/subagents/*.jsonl

# Codex CLI — "rollouts", dated folders
~/.codex/sessions/YYYY/MM/DD/rollout-<timestamp>-<uuid>.jsonl

On Windows the same trees live under %USERPROFILE%. Claude Code encodes the project’s absolute path into the directory name by replacing separators with dashes (/Users/you/dev/webapp-Users-you-dev-webapp). Subagent transcripts — what the Task tool spawns — are separate files beside the parent session in newer versions. Long-running histories are not small: individual session files reach tens or hundreds of megabytes.

What a record looks like

Each line is one JSON record. The load-bearing field is type; the common ones are user, assistant, system, and summary, and new types appear as Claude Code evolves. A trimmed assistant record:

{"type":"assistant","uuid":"…","parentUuid":"…","timestamp":"2026-07-01T10:00:05.123Z",
 "sessionId":"…","cwd":"/Users/you/dev/webapp","version":"1.0.61","isSidechain":false,
 "message":{"id":"msg_01…","model":"claude-opus-4-8","role":"assistant",
   "content":[{"type":"text","text":"I'll look at the hook first."}],
   "usage":{"input_tokens":4,"output_tokens":85,
            "cache_read_input_tokens":5496,"cache_creation_input_tokens":0}}}

The parts that matter when you parse:

Trap one: usage repeats per line

Because one API response is split across one line per content block, every one of those lines repeats the same message.id and the same usage object. Sum usage over lines and your token and cost numbers come out roughly 2.5–3× too high — a mistake with precedent: an early Turnlog release did exactly this, overcounted about 2.7×, and the fix (count once per message.id) is on its changelog. Dedupe first:

# tokens for one session, counted once per API response
jq -s '[.[] | select(.message.usage) | {id: .message.id, u: .message.usage}]
  | unique_by(.id) | map(.u.input_tokens + .u.output_tokens) | add' session.jsonl

Codex rollouts, and trap two

Codex CLI’s files are also JSONL but differently shaped. Two things to know before aggregating:

If you’re building on this

Three rules earn their keep: dedupe usage by response id (both traps are versions of this), classify records defensively and keep the unknowns raw, and treat the format as adversarial — version-sniff, don’t assume. The one thing you can rely on is that everything is there: prompts, diffs, shell output, token counts, timestamps. It’s a complete flight recorder with no cockpit.

Turnlog replaying a parsed session: speaker rails for you and Claude, Read/Edit/Write tool rows, a colored diff, a nested subagent run, and an unrecognized event kept raw
The same records, parsed: turns, tool rows, diffs — and an unrecognized event kept raw instead of dropped. (Sample sessions via npx turnlog demo.)

The parsed way

Turnlog reads all of this for you

Turnlog is a free, MIT-licensed CLI whose adapters handle both formats — usage deduped by response, resumes stitched into one conversation, Codex’s two channels and cumulative counters handled, unknown records kept raw and counted on a health panel instead of crashing. One command indexes everything into local full-text search and turn-by-turn replay. You can even drop a JSONL file in the browser demo and watch it parse — the page can’t transmit it.

Related: searching your history · tracking costs honestly · one file’s history across sessions