Concepts¶
The core idea: two kinds of sync¶
Everything scrubjay moves is one of two things — and which one it is decides the mechanism. Get this distinction and the rest of the system follows:
| Semantic | Meaning | Mechanism | What it fits |
|---|---|---|---|
| Shared / bidirectional ("cross-machine") | Same content on every machine; edits merge | git (pull + push) | things you author: CLAUDE.md, commands, agents, settings, plugins, memory |
| Archive / one-way | machine → NAS; never edited in two places, no read-back | rsync (the P2P "cart") | records: transcripts, subagents, plans, readable/, history.jsonl, tasks |
The second axis is privacy, and it's orthogonal: anything sensitive goes straight
to your own NAS, never a third party. So the records ride peer-to-peer rsync to the
NAS, and the one piece of authored content that's sensitive — memory (it carries
real file paths) — still uses git for the merge, but a git repo self-hosted on the NAS
over WireGuard rather than GitHub. Only the non-sensitive authored config rides GitHub
(scrubjay-data). That's the whole design in one sentence: author-vs-record picks
git-vs-rsync; sensitive-vs-not picks NAS-vs-GitHub.
NAS or GitHub — your choice of shared store¶
The record half above is where a NAS shines, but a NAS isn't required. The transcript transport is pluggable, and the two backends are genuinely parallel — you pick one when you onboard:
- Your own NAS (
rsync-wg/local) — records ride peer-to-peer to it over WireGuard/SSH; nothing ever touches a third party. The tradeoff is standing up a NAS + WireGuard. - GitHub (
git) — each session is pushed to a privatescrubjay-chatsrepo. Zero infrastructure to run; the tradeoff is that your transcripts live in a (private) third-party repo rather than only on your own hardware.
Same records, two destinations — choose by whether you'd rather manage your own storage or
none. (Config always rides GitHub either way. Cross-machine memory rides its own repo and
follows the same fork: the NAS backends self-host it on your hardware, while a GitHub-only
setup puts it in a separate private scrubjay-memory repo — wired for you by /sjmemory, with
the same "your file paths now sit with a third party" trade-off you accepted for transcripts.
See memory.)
Keeping records out of git doesn't mean giving up recovery: on the NAS backends, point-in-time
history comes from filesystem snapshots of scrubjay-storage, not git — cheap for append-only
data, and the reason transcripts don't belong in a repo. See durability.
What is scrubjay?¶
Claude Code reads its configuration from a ~/.claude/
directory on whatever machine you run it on: which rules it must follow, which custom
commands and sub-agents exist, what it's allowed to do, and so on. If you work on several
machines (a laptop, a desktop, an HPC cluster) you'd normally set all of that up by hand,
separately, on each one — and they'd drift apart over time.
scrubjay keeps that configuration in git instead, and makes it apply itself. This
repo (scrubjay) holds only the machinery — the shell scripts and hooks. Your actual
content lives in a second, private repo (scrubjay-data) that we call the database.
A small sync step turns the database into a working ~/.claude/ on each machine, and a
pair of hooks keep it current automatically. The result: you configure Claude once, and
every machine stays in step.
What's in the database (scrubjay-data) and how Claude uses it¶
The database is just plain Markdown and JSON files, organised into a handful of directories. Each one feeds Claude in a specific way:
-
claude-md/— the configuration shared by all your machines.CLAUDE.mdis the always-on instruction file Claude reads at the top of every session (e.g. "never add aCo-Authored-Bytrailer to commits").commands/holds custom slash-commands you can invoke by name —commands/explain-diff.mdbecomes/explain-diff, which might tell Claude to summarise your staged git changes. (The generic/sj*commands aren't here — they ship with the app;claude-sync.shmerges both into~/.claude/commands/.)agents/holds sub-agents Claude can delegate to —agents/test-runner.mddefines a focused helper that runs your test suite and reports back.CLAUDE.mdandagents/are symlinked straight into~/.claude/, so editing a file here changes Claude's behaviour everywhere on the next pull. -
hosts/<machine>/— the part that is different per machine.env.mddescribes that box in prose (its OS, where Python lives, cluster quirks) so Claude knows the lay of the land;claude/settings.jsonholds machine-specific overrides (for example, this HPC node auto-accepts edits);chats.index.jsonis an auto-generated catalogue of which chats live on that machine. Because hosts sit at the top level, one machine's setup never leaks into another's — and you can ask Claude to "readhosts/laptop/and adapt it for this HPC box". -
settings/—settings.base.jsonis the baseline~/.claude/settings.jsonthat applies everywhere: the permission allow/deny lists (which shell commands Claude may run without asking), the default model, and which hooks fire. The sync step merges this baseline with the per-host overrides fromhosts/<machine>/into the final settings file. -
memory/— durable facts Claude has learned and should remember across sessions, one fact per file. Claude Code's built-in auto-memory lives per-project at~/.claude/projects/<project>/memory/;claude-sync.shsymlinks each of those into this repo undermemory/<host>/<project>/, so auto-saved memories are synced and sorted by the machine they were created on, then by project (mirroring Claude's native layout). -
templates/— reusable starting points for project-level config, kept out of the always-on path. A file liketemplates/<project>/CLAUDE.local.mdis a ready-made rules file you can drop into a specific project so Claude picks up that project's conventions. -
logs/— a human-readable history of every session, one file per machine (<host>.log). Each line records the time, machine, working directory, and your first prompt, so you can latergrepfor "that chat about the auth refactor" across all machines. -
runbooks/— operational notes for you (not Claude), like the plan for moving chat transcripts off GitHub onto a private WireGuard link.
How the system works, and what you set up once¶
The moving parts fit together like this:
-
Two repos, one machine-local pointer. You clone
scrubjay(this machinery) andscrubjay-data(your content) onto each machine. A tiny file at~/.config/scrubjay/configtells the scripts where those clones live, and~/.config/scrubjay/hostpins a stable name for the machine (handy on clusters whose hostnames change between logins). -
Sync turns the database into
~/.claude/. Runningbin/claude-sync.shsymlinks the sharedclaude-md/scopes into~/.claude/, points each project's auto-memory dir at the syncedmemory/<host>/<project>/, and mergessettings/+ the host overrides into~/.claude/settings.json. It's safe to re-run; it only changes what actually differs. -
Two hooks keep it hands-off. When a session starts, a hook pulls the latest
scrubjay-dataand re-runs sync, so config you edited on another machine is already in effect. When a session ends, a hook appends the log line, refreshes that machine's chat index, and commits both back. You don't run these — they run themselves.
What you actually do once per machine: clone the two repos, write the little config
pointer, run claude-register-host.sh to scaffold a hosts/<machine>/ entry, then
claude-sync.sh to apply it. After that, you only ever edit files in scrubjay-data and
push — every machine converges on its own. The concrete commands are in
Onboarding.
Cross-machine tailoring with Claude¶
Everything is plain Markdown/JSON, so from any project you can ask Claude: "read
scrubjay-data/hosts/<other>/ and adapt its rules for this box" — it reads one host's
config and writes another's.