Concepts¶
The core idea: two kinds of sync¶
Everything scrubjay moves is one of two things — and which one it is decides the mechanism. Get this distinction and the rest of the system follows:
| Semantic | Meaning | Mechanism | What it fits |
|---|---|---|---|
| Shared / bidirectional ("cross-machine") | Same content on every machine; edits merge | git (pull + push) | things you author: CLAUDE.md, commands, agents, settings, plugins, memory |
| Archive / one-way | machine → NAS; never edited in two places, no read-back | rsync (the P2P "cart") | records: transcripts, subagents, plans, readable/, history.jsonl, tasks |
The second axis is privacy, and it's orthogonal: anything sensitive goes straight
to your own NAS, never a third party. So the records ride peer-to-peer rsync to the
NAS, and the one piece of authored content that's sensitive — memory (it carries
real file paths) — still uses git for the merge, but a git repo self-hosted on the NAS
over WireGuard rather than GitHub. Only the non-sensitive authored config rides GitHub
(scrubjay-data). That's the whole design in one sentence: author-vs-record picks
git-vs-rsync; sensitive-vs-not picks NAS-vs-GitHub.
A third axis: what gets loaded¶
Among the authored content there is one more distinction, and it decides where a thing goes inside the memory repo rather than which repo it goes to:
| memory | note | |
|---|---|---|
| Shape | one fact per file | a document |
| Loaded | MEMORY.md is read at the start of every session |
never, until asked for |
| Written by | the agent, as it learns | you, via /sjnote |
| Lives in | <memory>/<project>/ |
<memory>/<project>/notes/ |
Both are authored and both are sensitive, so both take the same route — git, self-hosted. What
separates them is cost. MEMORY.md is budgeted (the harness loads its first 200 lines) because it
is paid for in every future session; a two-page analysis kept there is a permanent tax. A note is
the same durability with none of that: retrieved on demand through
/sjrecall or sj_list(type="note"), invisible otherwise.
This is also why notes are not archive records. A note is the one authored thing you may well want to edit six months later, and editing in two places needs a merge — which is exactly what the one-way archive gives up by design.
NAS or GitHub — your choice of shared store¶
The record half above is where a NAS shines, but a NAS isn't required. The transcript transport is pluggable, and the two backends are genuinely parallel — you pick one when you onboard:
- Your own NAS (
rsync-wg/local) — records ride peer-to-peer to it over WireGuard/SSH; nothing ever touches a third party. The tradeoff is standing up a NAS + WireGuard. - GitHub (
git) — each session is pushed to a privatescrubjay-chatsrepo. Zero infrastructure to run; the tradeoff is that your transcripts live in a (private) third-party repo rather than only on your own hardware.
Same records, two destinations — choose by whether you'd rather manage your own storage or
none. (Config always rides GitHub either way. Cross-machine memory rides its own repo and
follows the same fork: the NAS backends self-host it on your hardware, while a GitHub-only
setup puts it in a separate private scrubjay-memory repo — wired for you by /sjmemory, with
the same "your file paths now sit with a third party" trade-off you accepted for transcripts.
See memory.)
Keeping records out of git doesn't mean giving up recovery: on the NAS backends, point-in-time
history comes from filesystem snapshots of scrubjay-storage, not git — cheap for append-only
data, and the reason transcripts don't belong in a repo. See durability.
What is scrubjay?¶
Claude Code reads its configuration from a ~/.claude/
directory on whatever machine you run it on: which rules it must follow, which custom
commands and sub-agents exist, what it's allowed to do, and so on. If you work on several
machines (a laptop, a desktop, an HPC cluster) you'd normally set all of that up by hand,
separately, on each one — and they'd drift apart over time.
scrubjay keeps that configuration in git instead, and makes it apply itself. This
repo (scrubjay) holds only the machinery — the shell scripts and hooks. Your actual
content lives in a second, private repo (scrubjay-data) that we call the database.
A small sync step turns the database into a working ~/.claude/ on each machine, and a
pair of hooks keep it current automatically. The result: you configure Claude once, and
every machine stays in step.
What's in the database (scrubjay-data) and how Claude uses it¶
The database is just plain Markdown and JSON files, organised into a handful of directories. Each one feeds Claude in a specific way:
-
claude-md/— the configuration shared by all your machines.CLAUDE.mdis the always-on instruction file Claude reads at the top of every session (e.g. "never add aCo-Authored-Bytrailer to commits").commands/holds custom slash-commands you can invoke by name —commands/explain-diff.mdbecomes/explain-diff, which might tell Claude to summarise your staged git changes. (The generic/sj*commands aren't here — they ship with the app;claude-sync.shmerges both into~/.claude/commands/.)agents/holds sub-agents Claude can delegate to —agents/test-runner.mddefines a focused helper that runs your test suite and reports back.CLAUDE.mdandagents/are symlinked straight into~/.claude/, so editing a file here changes Claude's behaviour everywhere on the next pull.skills/holds skills —skills/<name>/SKILL.mdplus whatever that skill bundles — linked per directory into~/.claude/skills/and into opencode's ownskills/, sinceSKILL.mdis a cross-tool standard both harnesses read.output-styles/follows the same per-file shape ascommands/. -
shared/AGENTS.md— instructions that hold whichever agent you are driving. opencode is pointed at it by absolute path (opencode.json→instructions[]); Claude Code is given it as a user-level rule (symlinked to~/.claude/rules/, which loads on every session in every project). That indirection is deliberate: Claude Code readsCLAUDE.md, notAGENTS.md, so a rule is the only user-scope mechanism that makes one shared file reach both harnesses without rewriting theCLAUDE.mdyou authored. -
hosts/<machine>/— the part that is different per machine.env.mddescribes that box in prose (its OS, where Python lives, cluster quirks) so Claude knows the lay of the land;claude/settings.jsonholds machine-specific overrides (for example, this HPC node auto-accepts edits);chats.index.jsonis an auto-generated catalogue of which chats live on that machine. Because hosts sit at the top level, one machine's setup never leaks into another's — and you can ask Claude to "readhosts/laptop/and adapt it for this HPC box". -
settings/—settings.base.jsonis the baseline~/.claude/settings.jsonthat applies everywhere: the permission allow/deny lists (which shell commands Claude may run without asking), the default model, and which hooks fire. The sync step merges this baseline with the per-host overrides fromhosts/<machine>/into the final settings file. -
memory/— durable facts Claude has learned and should remember across sessions, one fact per file. Claude Code's built-in auto-memory lives per-project at~/.claude/projects/<project>/memory/;claude-sync.shsymlinks each of those into this repo undermemory/<host>/<project>/, so auto-saved memories are synced and sorted by the machine they were created on, then by project (mirroring Claude's native layout). -
templates/— reusable starting points for project-level config, kept out of the always-on path. A file liketemplates/<project>/CLAUDE.local.mdis a ready-made rules file you can drop into a specific project so Claude picks up that project's conventions. -
logs/— a human-readable history of every session, one file per machine (<host>.log). Each line records the time, machine, working directory, and your first prompt, so you can latergrepfor "that chat about the auth refactor" across all machines. -
runbooks/— operational notes for you (not Claude), like the plan for moving chat transcripts off GitHub onto a private WireGuard link.
How the system works, and what you set up once¶
The moving parts fit together like this:
-
Two repos, one machine-local pointer. You clone
scrubjay(this machinery) andscrubjay-data(your content) onto each machine. A tiny file at~/.config/scrubjay/configtells the scripts where those clones live, and~/.config/scrubjay/hostpins a stable name for the machine (handy on clusters whose hostnames change between logins). -
Sync turns the database into
~/.claude/. Runningbin/claude-sync.shsymlinks the sharedclaude-md/scopes into~/.claude/, points each project's auto-memory dir at the syncedmemory/<host>/<project>/, and mergessettings/+ the host overrides into~/.claude/settings.json. It's safe to re-run; it only changes what actually differs. -
Two hooks keep it hands-off. When a session starts, a hook pulls the latest
scrubjay-dataand re-runs sync, so config you edited on another machine is already in effect. When a session ends, a hook appends the log line, refreshes that machine's chat index, and commits both back. You don't run these — they run themselves.
What you actually do once per machine: clone the two repos, write the little config
pointer, run claude-register-host.sh to scaffold a hosts/<machine>/ entry, then
claude-sync.sh to apply it. After that, you only ever edit files in scrubjay-data and
push — every machine converges on its own. The concrete commands are in
Onboarding.
Cross-machine tailoring with Claude¶
Everything is plain Markdown/JSON, so from any project you can ask Claude: "read
scrubjay-data/hosts/<other>/ and adapt its rules for this box" — it reads one host's
config and writes another's.