The knowledge base

Verified findings about AI coding harnesses — Claude Code and its neighbours — written down so the next person does not pay for them again.

How this differs from docs/lessons.md

Both hold things learned the expensive way. The line between them is the audience:

  • docs/lessons.md — what this plugin learned. Decisions and failures that shaped floppy’s own code. A reader who does not use floppy has no use for them.
  • here — what is true regardless of floppy: behaviour of the harness, traps in shell and git, how agent memory rots. A reader who does not use floppy is exactly the audience.

When a lesson turns out to hold outside this plugin it moves here, and lessons.md stops carrying it. Two did, on 2026-09-05.

What belongs here

A note is admitted only if all three hold:

  1. True outside one repository. A fact about your own codebase belongs in that codebase’s agent memory, not here.
  2. Cost real work to learn. Reading it in the vendor documentation does not count. Finding it by disassembling a binary, by a failed run, or by a measurement does.
  3. Checkable. The note carries a date, the version or environment it was verified against, and a concrete way to re-verify it. A claim nobody can re-test is folklore.

Anything failing one of the three is not wrong, it is somebody else’s file. Advice (“delegate mechanical work to a cheaper model”) fails condition 2: anyone derives it in a week. A measurement of what the harness does, and how you confirmed it, does not.

Why the checkable clause is load-bearing

A project’s own memory is exercised every session. A stale note there gets in the way, somebody notices, somebody fixes it. A shared knowledge base has no such pressure. Nobody re-reads a note about somebody else’s tool until they are already burned by it, so a wrong entry rots quietly and then actively misleads. Bases like this do not usually die of emptiness; they die of confident, out-of-date entries.

So every note carries verified_on and recheck, and

python3 scripts/knowledge-rot-check.py

lists the ones that have aged out. It reports, it does not gate — old and wrong are different things, and only a person who knows the area can tell them apart.

Layout

knowledge/
  README.md          you are here — what belongs, and why
  CONTRIBUTING.md    how to write a note: the contract and its required fields
  LINKS.md           existing practice in this space, with an assessment of each
  _template.md       copy this to start a note
  notes/
    harness/         behaviour of the coding harness itself
    memory/          how agent memory is organised, selected, kept from rotting
    shell/           shell, git and OS traps that reproduce anywhere
    practice/        ways of working that are measured, not merely recommended

There is deliberately no index file. The page on the documentation site is generated from the notes’ own front matter by scripts/site-build.sh, the same way the skills page is generated from skills/*/SKILL.md; a hand-written index would be a second copy of it, and the two would differ within a month. On GitHub the directory listing is the index.

The notes

Each entry is generated from the note’s own front matter: what it claims, what it was verified against, and how to check it is still true.

harness

Under an agent’s tool runner stdin is a pipe that never reaches EOF, so any command that reads it hangs forever while CI stays green

  • verified 2026-08-25 against Claude Code tool runner on Linux; original measurement by the floppy author, not re-run since
  • recheck: Run a script that reads stdin without redirection through the agent's shell tool; it should not return. Then repeat with </dev/null.
  • confirmed by a person, not by a runnable check
  • the note on GitHub

A background session is refused on its first edit inside the session’s own repository — but not in additional directories — and the escape is a worktree branched from your own head, entered by path rather than by name

  • verified 2026-09-01 against Claude Code background job in a repository with a feature branch of its own; EnterWorktree with name: and with path:, and .claude/settings.local.json
  • recheck: Start a background session in a repository and have it edit one tracked file: the refusal names the missing isolation. Repeat the edit in a directory added as an additional directory — it succeeds.
  • confirmed by a person, not by a runnable check
  • the note on GitHub

The harness tracks files opened with Read and deduplicates them; a file pulled in with cat or sed through Bash arrives in full every single time, and editing one outside Edit re-injects the whole file

  • verified 2026-08-29 against Claude Code, one finished slice of work in a multi-repository project: 3 repositories, 7 commits, 4 review rounds; message content measured from the session transcript
  • recheck: Read the same file twice with the Read tool and then twice with cat through Bash, and compare the input token counts on the four turns
  • confirmed by a person, not by a runnable check
  • the note on GitHub

After a compaction boundary only the stable prefix — 15–31k tokens of system prompt, tools and rules — is read from cache; everything else is written again, so delaying a compact to protect the cache protects nothing

  • verified 2026-09-08 against Claude Code, 12 compaction boundaries across every project on one machine; message.usage read on the turns either side of each subtype: compact_boundary record
  • recheck: grep the session transcripts for records with subtype compact_boundary, then compare cache_read_input_tokens on the turn before and the turn after each one
  • confirmed by a person, not by a runnable check
  • the note on GitHub

The “do not use the Agent tool” line is a built-in default of the Opus 5 prompt bundle, not anything in your config

  • verified 2026-09-16 against Claude Code 2.1.267 (native binary), Linux 6.18 (WSL2), model claude-opus-5
  • recheck: grep -ac 'tool, workflows, or deep-research unless the user' \"$(command -v claude)\"
  • checked by machine on any platform — scripts/knowledge-recheck.py
  • the note on GitHub

Claude Code’s Prometheus exporter binds dual-stack *:9464, which a container cannot reach even though curl from the host answers — OTEL_EXPORTER_PROMETHEUS_HOST=0.0.0.0 is required, and the symptom points at the network instead

  • verified 2026-09-05 against Claude Code 2.1.232 on WSL2 (net.ipv6.bindv6only=0), Docker Desktop 29.7.2, prom/prometheus:latest with –add-host=host.docker.internal:host-gateway
  • recheck: Start a session with OTEL_METRICS_EXPORTER=prometheus and no host override, then ss -ltn grep 9464 (shows *:9464) and check the Prometheus target: down with connection refused while curl localhost:9464/metrics on the host answers
  • confirmed by a person, not by a runnable check
  • the note on GitHub

Subagents of one type write to one memory directory, so two lenses launched together overwrite each other’s note — and the survivor may report a finding as a duplicate of something the reader never saw

  • verified 2026-09-06 against Claude Code, two subagents of the same reviewer type launched in one message, writing under .claude/agent-memory//
  • recheck: Launch two subagents of one type in the same message, each told to write a memory note about its own findings, then list .claude/agent-memory/<type>/ and compare against what each reported
  • confirmed by a person, not by a runnable check
  • the note on GitHub

A command that moves into a plugin is recorded in the transcript as /plugin:name, so any tool searching for the bare /name silently measures only the pre-plugin era while still printing a full table of history

  • verified 2026-09-05 against Claude Code 2.1.232, transcripts under ~/.claude/projects//*.jsonl, a repository whose /start and /wrap moved into a plugin
  • recheck: grep -oh '<command-name>[^<]*' ~/.claude/projects/*/*.jsonl | sort | uniq -c — runs from before the move appear as /name, runs after as /plugin:name
  • confirmed by a person, not by a runnable check
  • the note on GitHub

Claude Code’s Prometheus endpoint does carry tokens and cost, but a metric family only appears once it has been recorded, and tokens are recorded on the session’s second API request — so a one-request session serves the session counter alone, however long it lives

  • verified 2026-09-05 against Claude Code 2.1.232, WSL2, OTEL_METRICS_EXPORTER=prometheus, three headless controls on separate ports plus one live interactive session
  • recheck: Run a session with CLAUDE_CODE_ENABLE_TELEMETRY=1 OTEL_METRICS_EXPORTER=prometheus, make it issue at least two API requests (any tool call does), and curl localhost:9464/metrics grep ‘^# TYPE’ while it is alive
  • confirmed by a person, not by a runnable check
  • the note on GitHub

Checkpoints cover only file-editing-tool edits — bash changes, background subagent edits and symlinked paths survive a rewind

  • verified 2026-09-05 against Claude Code 2.1.232; docs at code.claude.com/docs/en/checkpointing
  • recheck: Read code.claude.com/docs/en/checkpointing, section 'Limitations'; confirm the four exclusions are still listed
  • confirmed by a person, not by a runnable check
  • the note on GitHub

A subagent’s first turn reads zero tokens from cache — it cannot use the parent’s prefix and pays to build its own, so delegation only pays back above that entry price

  • verified 2026-08-21 against Claude Code, one general-purpose subagent on Sonnet, four turns; usage fields read per turn from the subagent transcript
  • recheck: Run any subagent, open ~/.claude/projects/<checkout>/<session>/subagents/agent-<id>.jsonl, and read message.usage on the first assistant turn: cache_read_input_tokens is 0 and cache_creation_input_tokens is the whole prefix
  • confirmed by a person, not by a runnable check
  • the note on GitHub

Adding memory: to a subagent grants it Write and Edit over the whole workspace, whatever its tools: list says

  • verified 2026-09-05 against Claude Code 2.1.232, Linux 6.18 (WSL2), subagent with memory: project
  • recheck: Give a subagent tools: Read, Bash plus memory: project, ask it to Write a file outside its memory directory, and compare md5 before and after
  • confirmed by a person, not by a runnable check
  • the note on GitHub

memory

A memory linter that checks form reports a corpus clean while fifteen of its notes state something that has stopped being true, and an existence check does not close the gap

  • verified 2026-09-05 against floppy 0.15.0 memory linter; a 157-note, 489 012-character agent-memory corpus grown over five weeks by two machines
  • recheck: Take a memory corpus that has run for a month, run the linter (expect clean), then read every note against the current repository and count the notes whose claims no longer hold
  • confirmed by a person, not by a runnable check
  • the note on GitHub

shell

A false [[ cond ]] && cmd as the last line makes the whole script exit 1, so hook runners report failure on success

  • verified 2026-09-05 against bash 5.x on Linux 6.18 (WSL2); behaviour is POSIX and holds on bash 3.2 / macOS
  • recheck: printf '%s\\n' '#!/usr/bin/env bash' 'f=1' '[[ $f -eq 0 ]] && echo x' > /tmp/r.sh; bash /tmp/r.sh; echo $?
  • checked by machine on linux, macos — scripts/knowledge-recheck.py
  • the note on GitHub

git check-ignore exits 128 for a path that traverses a symlink — a refusal to answer, which the two-valued idiom if ! git check-ignore -q reads as “not ignored” and acts on

  • verified 2026-09-08 against git on Linux 6.18 (WSL2); measured in this plugin while wiring a scope whose container sits under a symlinked directory
  • recheck: In a repository, symlink a directory and ask about a path underneath the link: git check-ignore -q -- link/file, then echo $?
  • checked by machine on linux, macos — scripts/knowledge-recheck.py
  • the note on GitHub

find over a symlinked directory returns nothing and exits 0, so a checker that walks one passes while looking at zero files

  • verified 2026-09-05 against GNU findutils on Linux 6.18 (WSL2); POSIX behaviour, holds on macOS
  • recheck: mkdir -p /tmp/r/real && touch /tmp/r/real/{a,b}.md && ln -s /tmp/r/real /tmp/r/link && find /tmp/r/link -name '*.md' | wc -l
  • checked by machine on linux, macos — scripts/knowledge-recheck.py
  • the note on GitHub

The tool shell is initialised from the user profile, so cp/mv/rm may be aliased to -i; with no terminal on stdin the prompt goes unanswered and the file is not overwritten, while the exit code says either 0 or 1 depending on the coreutils version

  • verified 2026-09-06 against Claude Code Bash tool on Linux 6.18 (WSL2), profile aliasing cp to cp -i; GNU coreutils 8.32 locally and a GitHub-hosted ubuntu runner
  • recheck: type cp — then run cp -i over an existing file with </dev/null and read the destination: it still holds the old content. Check the exit code too, and expect it to differ between coreutils versions
  • checked by machine on linux — scripts/knowledge-recheck.py
  • the note on GitHub

A green macOS job proves nothing about bash 3.2 compatibility — bash in PATH is 5.x, and calling /bin/bash run.sh still hands every test back to 5.x unless the runner passes the same interpreter down

  • verified 2026-08-25 against macOS on arm64 (darwin25); /bin/bash 3.2.57(1)-release against 5.x from Homebrew earlier in PATH
  • recheck: On a mac: bash –version head -1 and /bin/bash –version head -1 — the first says 5.x, the second 3.2.57. Then run your test suite under /bin/bash and print the interpreter from inside one test.
  • checked by machine on macos — scripts/knowledge-recheck.py
  • the note on GitHub

**A macOS runner’s temp directory is /var/folders///T/, and is not alphanumeric — it can contain _. Measured, it is fixed by the runner IMAGE, not drawn per machine, so while the pool serves two images the same commit passes on one and fails on the other and the split reads as flakiness**

  • verified 2026-09-06 against GitHub-hosted macos-latest (arm64); the split measured 2026-09-06 over 20 runners in one dispatch, run 34028462447 of spscream/ai-floppy; first observed 2026-09-05 in runs 33988970539 and 33989864363
  • recheck: On a mac or a macOS runner: d=$(mktemp -d); echo $d — expect /var/folders/<a>/<b>/T/tmp.XXXXXXXX. For the underscore and its frequency, dispatch .github/workflows/tmpdir-probe.yml, which runs knowledge/probes/tmpdir-probe.sh on 20 runners at once and tallies the components.
  • checked by machine on macos — scripts/knowledge-recheck.py
  • the note on GitHub

A [a-z] range in a shell pattern is matched through LC_COLLATE and on macOS it catches uppercase, while [a-z] in a Python regex is a codepoint range no locale touches — so the same rule written in both languages agrees on Linux under every locale and disagrees on macOS

  • verified 2026-09-06 against bash 5.1.16 and glibc 2.35 on Linux 6.18 (WSL2), under C.UTF-8 and a localedef-generated en_US.UTF-8; Python 3.12.3; the macOS half on GitHub macos-latest, job macos-bash-3-2 of spscream/ai-floppy — as a test failure under /bin/bash 3.2.57 in PR #36, and under the pinned locale in run 34060322874
  • recheck: LC_ALL=en_US.UTF-8 bash -c 'case RU in [a-z][a-z]) echo matches;; *) echo skips;; esac' — prints matches on macOS, skips on glibc. Compare with python3 -c 'import re; print(bool(re.match(r\"^[a-z]{2}$\", \"RU\")))', which prints False everywhere.
  • checked by machine on macos — scripts/knowledge-recheck.py
  • the note on GitHub

Under zsh an unquoted –include=*.ext aborts the command with “no matches found” and $b:path is read as a history modifier, so the measurement returns nothing rather than something wrong

  • verified 2026-09-08 against zsh 5.9 on Linux 6.18 (WSL2), commands issued through a tool whose shell is zsh
  • recheck: Run zsh -c 'b=Feature; print -- ${b:l}' — it prints feature, proving :l is a modifier — and run grep -r --include=*.ext in a directory with no such file at the top level
  • checked by machine on linux, macos — scripts/knowledge-recheck.py
  • the note on GitHub

practice

a compound cd A && … ; cd B && … in one tool call does not measure two trees — the working directory resets between calls and the second cd may not run at all, so both halves report the same tree

  • verified 2026-09-08 against Claude Code Bash tool, two repositories side by side; paid four times, twice on the day the note was written
  • recheck: Run cd A && pwd ; cd B && pwd in one tool call, then run the same two commands as two separate calls, and compare the four paths
  • confirmed by a person, not by a runnable check
  • the note on GitHub

A check that cannot establish its own condition has to fail closed — that default is never written as a decision, and when it points the wrong way the broken guard reports “clean”

  • verified 2026-09-05 against floppy 0.15.1 — scripts/wrap-lock.sh, scripts/wrap-guard.sh, shim/run; the original failure on macOS arm64 under /bin/bash 3.2.57
  • recheck: Give a guard an input on which it cannot compute its own condition — a stat spelling the platform lacks, an unreadable file, an empty config key — and read the verdict, not the exit code. It should refuse.
  • confirmed by a person, not by a runnable check
  • the note on GitHub

lunr indexes every non-Latin word as the empty string, because its trimmer strips \W from both ends of a token and JavaScript’s \w is ASCII-only — the tokenizer is not involved, so tokenizer_separator is the wrong place to look

  • verified 2026-09-06 against lunr 2.3.9 as shipped by just-the-docs 0.12.0; measured against a deployed site with Russian pages; Node 18.14.2
  • recheck: In a page with lunr loaded: lunr.trimmer(new lunr.Token(‘память’)).toString() returns ''. Against a whole site, rebuild its index from the served search-data.json and count how many terms of Object.keys(idx.invertedIndex) contain non-Latin characters.
  • checked by machine on any platform — scripts/knowledge-recheck.py
  • the note on GitHub

**A template that serves HTML as a single line deletes every newline inside an inline

  • verified 2026-09-06 against Jekyll 4.4.1 with just-the-docs 0.12.0 on GitHub Pages; the JavaScript half is language behaviour, checked on Node 18.14.2
  • recheck: Fetch a page the template serves and count newlines inside its inline <script>: curl -s | node -e 'let h=\"\";process.stdin.on(\"data\",d=>h+=d).on(\"end\",()=>{const m=h.match(/
  • checked by machine on any platform — scripts/knowledge-recheck.py
  • the note on GitHub

After a fetch, git diff main..branch shows everything other people added as deletions by your branch; what a branch contributes is the three-dot diff, counted from the merge base

  • verified 2026-09-08 against git on Linux 6.18 (WSL2); reproduced in a scratch repository with one commit on each side
  • recheck: In a scratch repository commit one file on a branch and a different file on the base, then compare git diff --name-only base..branch with base...branch
  • checked by machine on linux, macos — scripts/knowledge-recheck.py
  • the note on GitHub

Back to top

MIT licensed. The pages of this site are generated from the repository's own documents by scripts/site-build.sh — edit those, never a page.

This site uses Just the Docs, a documentation theme for Jekyll.