Architecture decisions
Short record of the choices made and why, so future-me doesn’t relitigate them.
D1 — Build on Herdr’s socket API, don’t rebuild a multiplexer
Section titled “D1 — Build on Herdr’s socket API, don’t rebuild a multiplexer”Herdr exposes herdr api snapshot (full state as JSON), agent/pane
read+control, and agent wait (blocks until a state change). That’s a complete
read + control + event surface. gothalo consumes it; it invents nothing Herdr
already does.
D2 — A bridge daemon is required (push forces it)
Section titled “D2 — A bridge daemon is required (push forces it)”A phone can’t run herdr agent wait. Something long-lived next to Herdr must
watch state and reach out to FCM. That daemon is the bridge. Consequence: a
“phone SSHes in and runs herdr” design (Moshi’s model) is rejected — SSH is
pull-only and can’t push.
D3 — Push is outbound; interactive is tailnet-only
Section titled “D3 — Push is outbound; interactive is tailnet-only”FCM = bridge → Firebase (outbound), so notifications need zero inbound exposure and work on any network. Only the interactive API (snapshot/type/ terminal) needs the phone to reach the bridge, and that goes over the tailnet. Two separate concerns; don’t conflate their networking.
D4 — Tailscale for exposure, not raw ports or SSH tunnels
Section titled “D4 — Tailscale for exposure, not raw ports or SSH tunnels”Already running Tailscale. Bridge binds to the tailnet IP; the tailnet is the auth boundary; a per-user bearer token allows revocation. No public ports, no port-forwarding, no SSH client on the phone. (An SSH tunnel would replace Tailscale as the exposure layer, not sit alongside it — not chosen.)
D5 — Flutter over React Native
Section titled “D5 — Flutter over React Native”The core screen is a fast-scrolling terminal. Flutter’s xterm.dart is a
native terminal widget; React Native has no native terminal and would embed
xterm.js in a WebView (JS-bridge boundary on every keystroke/output chunk —
wrong seam for high-frequency streaming). RN’s only edge (team already writes
React/TS) doesn’t outweigh the terminal being the product. Revisit only if v1
becomes “status board first, terminal later.”
D6 — The mobile keyboard is an accessory row, not a custom IME
Section titled “D6 — The mobile keyboard is an accessory row, not a custom IME”Soft keyboards lack Esc/Ctrl/Tab/arrows. Solution is a toolbar above the system
keyboard whose buttons write control bytes into the same PTY stream (Esc=0x1b,
Ctrl+C=0x03, arrows=\e[A…). A sticky Ctrl toggle (letter & 0x1f) collapses
the whole Ctrl-combo space into one button — it now lives in the more-sheet
below, but the mechanism is unchanged. ~15 lines of UI, not a keyboard
extension.
Arrows moved out of that row into an arrow pad: ↑ over ← ↓ → over ⌫ beside a double-width Enter, with hold-to-repeat on everything except Enter (a leaned-on thumb must not resubmit). Arrows are what agent TUIs (menus, history, approval prompts) ask for most, but four more keys in a scrolling row is a poor way to offer them.
A popup, not a fixture: it costs nothing while closed, and one tap dismisses it when it’s over something you want to read. An earlier draft floated permanently over the buffer and was draggable to get it out of the way — dragging is a worse answer to “it’s covering something” than closing is, and having arrows in both the pad and the row meant two homes for one key.
The row itself is one strip of exactly seven small equal buttons spread evenly, and — this is the part that took two goes to get right — the same seven every time:
Esc ^C ⋯ [pad] Tab ⌨ 🖼Even spacing is what makes it read as one control surface rather than a huddle of chips, and seven puts the pad toggle on the exact centre line, directly under the pad it opens. The keyboard toggle lives here rather than in the pad, both for that count and because it’s a screen control, not a keystroke; without it the soft keyboard only ever appeared as a side effect of tapping the buffer, which is also how you scroll it. ^C is the one control byte with a place of its own: it’s the emergency stop and must never cost two taps.
Everything below the buffer is one AccessoryButton — same fill, radius, height
and mono type, width the only variable. Two earlier attempts are worth not
repeating: the quick commands as Material chips (outlined, proportional,
stadium) stacked above filled mono key blocks read as two unrelated toolbars;
and packing labelled chips plus a pinned toggle plus four keys into one row
needs ~470dp of a ~393dp phone, so something always clipped mid-word.
The row is fixed because a config-shaped row has no centre
Section titled “The row is fixed because a config-shaped row has no centre”An earlier version of this row was seven by default and grew from there: it
prepended one button per saved quick command and carried + (add one) and a
sticky Ctrl. That made its width, and therefore its centre, a function of the
user’s config. One saved command was enough to push the pad toggle off the
centre line; adding the image button pushed the strip past the screen entirely,
at which point spaceEvenly has no free space to distribute, the row
left-aligns and scrolls, and the toggle sits visibly off-centre under the pad it
opens. Shipped, spotted on a Pixel-class phone within minutes, and correctly
called out: the earlier note here claimed appending “moves nothing”, which
confused preserving the order with preserving the centre. It preserved
neither for long.
So the variable half moved behind ⋯ into a more-sheet, and the row became
a constant:
- Control bytes —
^C ^D ^Z ^L ^R ^U ^W ^A ^E. Not every combination: the ones that earn a button on a phone. End or detach (^C ^D ^Z), redraw a garbled repaint (^L), search history instead of typing a long command (^R— arguably the most valuable key here), fix a typo without forty backspaces (^U ^W), and jump to line start/end (^A ^E), which a soft keyboard has no Home/End for. - Sticky Ctrl — kept, in the sheet. The nine bytes above are a shortlist,
not the space;
^K,^P/^N,^B/^F,^X,^G,^]all matter to somebody, and a phone keyboard has no Ctrl key of its own, so dropping this would make them unreachable rather than merely slower. It arms and closes the sheet, and the row’s⋯lights while armed — otherwise the armed state would be invisible and the next letter would come out mangled with no warning. - Quick commands and
+— all of them, none pinned in the row. Pinning even one would put the count back under the user’s control and the centring would drift again. The old “filter out a command that duplicates a row key” rule goes with them: it existed because a slot in the strip was scarce, and a sheet has room — quietly hiding something the user saved is the worse trade.
Every sheet action closes the sheet, because its result is on the terminal behind it; multi-key work is what the arrow pad is for.
Attaching an image is the seventh button. It belongs below the buffer rather
than up in the app bar for the same reason the composer’s paperclip sits inside
the input pill: it acts on what you are typing, not on the session. It types the
uploaded path into the PTY like any other keystrokes, with no carriage return —
same insert-don’t-send rule as the composer — so it works for whatever is
running in the pane, and works on a pane with no agent at all. The bridge
resolves that pane’s own cwd for the drop (see CONTRACT-image.md); a path is
just text, and nothing about typing one needs an agent to exist.
Seven buttons plus their gaps measure ~340dp against a 393–412dp phone, so the row fits with room to spare and the even spread is real rather than a scroll view’s left edge. The minimum gap is 4dp, not 6dp, precisely because 6dp put it ~2dp over on the narrowest common phone — and 2dp of overflow is all it takes to lose the property this row exists to hold. Tests pin the count, the toggle’s index, and that neither moves with the number of saved commands.
The transcript composer follows the same rules — the two screens are one
tap apart doing the same job, so a different toolbar vocabulary on each read as
an accident. Its actions row is the same evenly-spread AccessoryButtons
(mode, quick commands, +, jump, terminal); attaching an image moved inside
the composer pill, where every messaging app puts it and where it belongs, as
it acts on the message being written rather than on the session; and jump moved
down from the app bar, which is a stretch away at the top of a phone.
The composer’s row keeps its quick commands inline, and did not follow them into a sheet. Its width is config-shaped in the same way, but nothing there is anchored to its centre — no popup opens from it — so a wider row merely scrolls, which is the behaviour it was drawn for. And the chips are that row’s reason: the composer has no key strip, so a sheet would cost a tap on the surface where quick commands are used most and buy nothing. Same list, same store, same long-press-to-remove wherever it is shown.
The agent’s “working” state there is now a bar sweeping the seam above those
controls, with no thinking… label — the motion says it, and the word cost a
line of transcript. Pulsing dots were the wrong borrow: they promise an imminent
message, where this is a machine holding a turn open for anything up to ten
minutes.
D7 — Multi-agent is free
Section titled “D7 — Multi-agent is free”Herdr detects and normalizes ~20 agents below the API into one status model, so the app writes multi-agent UI once. The only per-agent code is an optional ~12- line keystroke map for one-tap approvals (each agent’s confirm key differs); fallback is “open the terminal and let the human type,” which needs no per-agent code at all.
D8 — Idempotent approvals
Section titled “D8 — Idempotent approvals”An approval push may sit on the lock screen for minutes while the agent’s state
moves on. Every approve action carries agent + state_change_seq (both in the
snapshot); the bridge no-ops if the agent isn’t still blocked at that seq. This
guard lives in the bridge so every approval surface (banner button, Live
Activity, in-app) inherits it.
D9 — Backend-first, app-last
Section titled “D9 — Backend-first, app-last”Validate snapshot → tailnet reach → notify trigger → real push (in a browser tab
via FCM web push) entirely with curl/browser before writing Flutter. The app is
drawn over a backend already trusted. See TESTING.md.
D10 — One gothalo binary: bridge + CLI (cobra)
Section titled “D10 — One gothalo binary: bridge + CLI (cobra)”The bridge is now a proper CLI (gothalo serve | pair | devices) built on cobra,
laid out as a standard Go module (cmd/gothalo + internal/*). pair/devices
are thin clients of the running serve daemon over a localhost admin API, so the
daemon stays the single source of truth for state.
D11 — QR pairing + per-device bearer tokens
Section titled “D11 — QR pairing + per-device bearer tokens”Devices onboard by scanning a QR (gothalo pair prints it) that encodes a
one-time code + a connect URL. POST /pair {code, device_name, fcm_token} issues
a per-device bearer, stored in ~/.gothalo/devices.json. This replaces the
single shared token and finally delivers real revocation (D4): gothalo devices revoke <id> kills one device’s bearer and stops its pushes without touching
the others. An operator admin token (auto-generated, in config.json) gates
the CLI/admin endpoints and the web test page.
D12 — Pluggable transport; Tailscale is one option, not a requirement
Section titled “D12 — Pluggable transport; Tailscale is one option, not a requirement”The HTTP API is a transport-agnostic http.Handler. A transport.Transport
seam runs it under direct mode (listen locally — behind tailscale serve,
or a LAN/tailnet IP) today, and a relay mode later: the bridge dials OUT to a
small hosted broker over a persistent WebSocket, so phones reach it through the
relay with no inbound ports and no Tailscale. Both feed the same handlers. Push
stays outbound (D3) and per-device bearers still gate access (D4) in either mode.
The pairing QR carries a generic connect endpoint so the app never hardcodes
Tailscale. Relay is stubbed now (internal/transport/relay), wired later.
D13 — Unified in-process event bus, streamed over WS /events
Section titled “D13 — Unified in-process event bus, streamed over WS /events”One process-wide Herdr subscription feeds an in-process pub/sub bus
(internal/events); every phone client is just another subscriber over WS /events. Fan-out is bounded — a subscriber that can’t keep up is dropped
(channel closed) and expected to reconnect and resync from the snapshot frame —
so one lagging phone can’t stall the bus or the others. This is the live-update
backbone the app builds on, and the seam an event/plugin ingestion layer plugs
into (see D19).
D14 — Multi-session bridge
Section titled “D14 — Multi-session bridge”herdr.Manager watches all Herdr sessions, not just the default, starting and
stopping per-session workers as sessions appear/disappear. People run more than
one Herdr session; the bridge must not be blind to the others.
D15 — Three views of an agent, not one
Section titled “D15 — Three views of an agent, not one”The app consumes an agent at three altitudes: the raw PTY (WS /attach, full
terminal), a parsed compact state card (agentstate — “what is it doing / what
is it asking”), and a normalized transcript chat (transcript). The phone
usually wants the semantic views; the PTY is the escape hatch / fallback.
D16 — Transcript from the agent’s own on-disk log, normalized
Section titled “D16 — Transcript from the agent’s own on-disk log, normalized”transcript tails the agent’s structured session file (Claude Code writes JSONL;
codex has its own format) and normalizes every entry — messages, thinking, tool
calls (command/diff), tool results — into one kind-agnostic chat schema, so the app
renders a single chat UI for any agent. Unrecognized entries pass through so the
tail survives schema drift. (Revisited in D19.)
D17 — Notification lifecycle via the bus
Section titled “D17 — Notification lifecycle via the bus”The notify-clearer (internal/notify), the bus’s first consumer, dismisses stale
“blocked” pushes: it remembers the pane behind each blocked push and, when the bus
shows that pane leaving blocked, clears the now-irrelevant notification — keeping
the lock screen honest (complements D8’s idempotent approvals).
D18 — /herdr allowlisted CLI proxy
Section titled “D18 — /herdr allowlisted CLI proxy”A single POST /herdr proxies an allowlisted set of Herdr operations
(worktree/tab/pane create+close, focus, …) so the app gets Herdr-parity controls
without a bespoke endpoint per verb, while the allowlist stops it from becoming an
arbitrary command sink.
D19 — Push-based plugin events supersede live file-tailing (direction)
Section titled “D19 — Push-based plugin events supersede live file-tailing (direction)”Each coding agent gains a gothalo plugin — its own native hook/plugin config
that POSTs normalized events (message, tool call, approval-needed, done) to
the bridge, which Publishes them to the event bus (D13). Push is real-time and
carries intent — “approval needed for Bash: rm -rf …” before the tool runs
— which tailing a transcript after the fact (D16) cannot give the notification /
approval path. It also drops the fragile per-agent file-path resolution.
Scope (deliberately not a hard delete of D16):
- Push becomes the primary live source; the app reads history + live from the bridge, not from agent files.
- File-reading is demoted, not removed: a one-time transcript import seeds pre-plugin history, and it stays the fallback for agents whose hook surface is too thin to reconstruct chat content.
- The normalized chat schema (D16) is the target every plugin maps into, so adding an agent = a plugin adapter, not a new file parser. Mirrors how Herdr normalizes status (D1/D7) — here gothalo normalizes events.
- Reality check: hook richness varies a lot — Claude Code is rich; opencode is event-native (client/server with an event stream); codex and others are thinner. Verify each agent’s real surface before writing its adapter. Where a plugin can’t carry full content, it fires on events and the bridge reads the transcript at that moment (hook-triggered) instead of continuously tailing.
Ingestion lands on a new path (POST /hook or similar) since GET /events is the
outbound stream. Spike with Claude Code first to prove the “approve with context”
UX, then generalize.
D20 — FCM credentials follow ADC; a shared key file is not the only path
Section titled “D20 — FCM credentials follow ADC; a shared key file is not the only path”internal/push accepts two credential shapes and finds them by Google’s
Application Default Credentials search order (configured path → $GOOGLE_APPLICATION_CREDENTIALS
→ gcloud’s well-known file → GCE/Cloud Run metadata server).
The motivation is team access, not flexibility. A downloaded service-account key
is a shared bearer secret: everyone holding the file is the same identity,
rotation breaks everyone at once, and the audit log can’t attribute a send. Adding
authorized_user support means a teammate runs gothalo push login, authenticates
as themselves via gcloud, and is granted/revoked individually in IAM — nobody
copies a key. This is also what makes the repo publishable: there is no shared
secret that must exist for a contributor to run the thing.
Consequences:
- The two shapes mint tokens by different grants (jwt-bearer vs refresh_token),
so
mintTokendispatches on shape. Scopes are bound at consent time for the refresh grant — hencepush loginpassing--scopesexplicitly, since gcloud’s default set omitsfirebase.messagingand the resulting 403 reads as a permissions bug rather than a scope bug. - User credentials name a person, not a project, so
push.project_idbecomes required on that path (push loginpersists it). push statusverifies with FCM’svalidate_onlyrather than sending, and maps 403 to a distinct error — “authenticated but never granted access” is the one failure a teammate cannot fix by logging in again.- The metadata-server branch means a future hosted relay (D12) can run with no key material anywhere. That is a free consequence of following ADC, not a commitment to build the relay.
D21 — Terminal scroll is a wheel report to the application, not scrollback
Section titled “D21 — Terminal scroll is a wheel report to the application, not scrollback”Dragging the live terminal scrolls the remote application, by sending it SGR mouse-wheel reports on the same PTY stream as every keystroke (D6). There is no client-side scrollback to scroll, and no host-side one either:
- Every agent pane runs on the alternate screen (herdr replays
?1049hon attach), so herdr keeps no scrollback for it —max_offset_from_bottomis0on every agent pane andpane read --source recentreturns exactly the visible frame. A bigger--linescannot recover what left the alt screen. - Herdr has no scroll-offset API (149 socket methods, none of them set
scroll), so the bridge can’t ask for a window of history either. - Claude Code turns on mouse tracking and SGR coordinates (
?1000h ?1002h ?1003h ?1006h), so it consumes wheel reports itself. Verified on a live pane:ESC[<64;20;20Mscrolls it.
Consequence, accepted: the pane’s own viewport moves, so a desktop operator watching that pane sees it scroll too. That is inherent to alt-screen apps — Moshi has it as well (its docs describe the same drag → wheel forwarding when attached to a multiplexer).
The shim is PtyMouseHandler: xterm.dart encodes wheel-up/down as buttons 68/69
(64 + 4, which sets the shift bit) instead of 64/65, and applications
ignore shift+wheel. Everything else about xterm’s gesture path already worked.
Not chosen — the two things the open-source herdr clients do instead, both of
which give up the live terminal: merino re-reads --source recent with a growing
line budget (400→2000) and renders it as text, which yields nothing on an
alt-screen agent pane; herdr-mobile-relay snapshots each pane every 4 s and
sequence-merges the diff into a reconstructed 10k-line history, which is lossy
and plain-text. For agent history gothalo already has the transcript (D16), read
from the agent’s own log — complete and structured. Plain (non-alt-screen) panes
are the one case where a recent read is worth having — see D22.
D22 — A plain pane’s scrollback is seeded once, not paged
Section titled “D22 — A plain pane’s scrollback is seeded once, not paged”A plain pane has no application to send a wheel to (D21) and no live byte stream:
the client receives whole frames prefixed with a clear-screen, so its buffer holds
nothing to scroll back through. WS /attach therefore sends one history frame
before the first repaint — pane read --source recent-unwrapped, without the
clear-screen prefix, so it lands in the emulator’s own scrollback. Later repaints
erase only the viewport (ED 2 leaves scrollback alone), so the seed survives.
Once, not paged, because Herdr caps pane read at 1000 rows and exposes no
offset parameter. Measured on two panes holding far more: a docker compose logs -f
pane with 10,467 rows and a server log with 1,960 — --lines of 1100, 1500 and
20000 all returned exactly 999. So 1000 rows is the entire reachable history and a
merino-style growing window buys nothing (merino’s own 2000-line cap is above what
Herdr will return). Deeper history needs an offset method Herdr does not have.
Unwrapped, because the captured rows are folded at the desktop’s column count;
replaying those folds on a phone double-wraps every long log line. Reading is
side-effect-free for the operator — on a pane sitting 1,254 rows back, a read left
offset_from_bottom untouched, so the old warning in bridge_client.dart (that
history could only be captured by physically scrolling the pane) no longer holds.
Accepted wart: the seed ends with the current frame, so a few lines can appear both in scrollback and on screen. Trimming by the pane’s row count would risk cutting past the overlap and leaving a silent gap, and a repeated line beats a lost one.
D23 — The Priority section is capped, but “needs you” is exempt
Section titled “D23 — The Priority section is capped, but “needs you” is exempt”Priority is automatic (every blocked and done agent lands in it) plus
whatever you starred, so with a dozen agents live it grew past the viewport and
pushed the servers list off the bottom of the home screen — the one screen that
is supposed to answer “what now” at a glance.
It collapses to five rows (kPriorityVisibleRows) with a Show N more
expander. Five two-line tiles plus the header and the expander leave two or
three server tiles visible on a ~390×780dp phone, which is the point: Priority
is the top of the home surface, not the whole of it.
The cap is soft in exactly one direction. An agent that needs you is never
behind the expander: the cut stretches down the list to cover the last blocked
row, so ten blocked agents render ten rows. More blocked agents than fit is
the case the app exists for, and a screen that hid them to stay tidy would be
tidy and wrong. Nothing stretches it the other way — done, working, idle
and starred rows all sit under the cap, so a starred idle agent can be behind
the expander.
Ordering is untouched. PriorityOverflow only ever cuts a prefix off the
list the bridge already ranked by attention_rank (see docs/API.md); it never
sorts, and the exemption is defined on the row (“does this one need
you”) rather than on its position, so an older bridge whose ranks arrive via the
local fallback still can’t bury a blocked agent. Starred does not promote a
row above the bridge’s rank — that would be a second, client-side priority
order, which is the thing attention_rank exists to prevent.
Collapsed, the footer carries a tally of the whole section — 3 need you · 5 done · 2 idle, in each status’s own badge colours — not of the hidden part.
A collapsed section should still say what it is sitting on; “5 hidden” alone
says only how much you are missing, never whether it matters. Expanded, the
counts go away (the rows say it) and only Show less remains — the control
outlives its own use, or the tap that opened the list would delete the only way
to shut it.
Expanded/collapsed is remembered for the session (a plain Notifier) and
shared by both surfaces that render the section — the home screen and the
Priority screen. It is one list drawn twice; open in one place and shut in the
other reads as a bug. A cold start comes back collapsed, which is the state that
fits the screen.
D24 — Deleting a worktree’s branch is the bridge’s own endpoint, and the safety rules live there
Section titled “D24 — Deleting a worktree’s branch is the bridge’s own endpoint, and the safety rules live there”Removing a worktree from the phone cleaned up the checkout and the workspace and left the branch behind, every time. That is not an oversight in the app: Herdr has no notion of branches anywhere on its socket, so there was nothing to proxy and nothing else in the system to pick it up. With worktree creation down to one tap, the refs pile up faster than anyone prunes them, and a phone is the one place with no way to prune.
So GET /branch-info + POST /branch-delete (see
CONTRACT-branch-delete.md) shell out to git
directly — the same exception /diff and internal/transcript already are, for
the same reason: the thing being read or done is not part of Herdr’s model.
They are deliberately not on the /herdr allowlist (D18). That proxy’s whole
design is “params verbatim, no server-side validation”, which is exactly wrong
for an operation whose entire substance is what it refuses to do. The rules —
never the default branch, whatever it is called; never a branch checked out in
any worktree; unmerged only on an explicit force — live in internal/gitbranch
and re-run on every delete, so they hold for any caller and never depend on
the preflight the client happens to be holding.
Two endpoints rather than one, because they answer either side of an operation
that can fail. The preflight needs the workspace to still exist (it is what
names the branch and the repo root); the delete needs the checkout to be gone
(git refuses a checked-out branch). Between them sits worktree.remove, and if
that fails the branch delete must not run — which only the caller can know.
Default off, and unmerged is a second dialog. Removing a worktree is
recoverable (worktree.open brings it back); deleting a branch is much less so,
and a phone is where a mis-tap is most likely. A merged branch is one checkbox.
An unmerged one names the commit count in its own confirm before the box will
tick — the same tap must not mean both things.
The default branch is resolved, never assumed. refs/remotes/<remote>/HEAD
first, then a conventional local name; if neither answers, nothing in that repo
is deletable. Assuming main is the single failure mode here with no recovery,
and a repo on trunk is not exotic.
Partial outcomes are reported as partial. “Worktree gone, branch kept” is a
normal result — unmerged, refused, or a bridge too old to have the endpoint —
and it says so rather than showing a generic success. Likewise the local delete
never touches the remote, so upstream and remote_deleted:false come back in
the payload and the UI states it. Pushing a branch deletion from a phone affects
everyone and reaches outside the host; nothing else on this bridge does that, and
this does not either.
D26 — deliberately unused
Section titled “D26 — deliberately unused”Held the one-tap “Create PR” decision while that feature was its own PR (#112). When #112 was absorbed into the pane-suggestions work, its reasoning moved into D29 — where it belongs, since “the agent performs this one” is a property of the suggestion mechanism rather than a decision standing on its own. The number is left vacant rather than recycled: renumbering would silently repoint every reference written while #112 was open.
D27 — Opening a space from the phone needs a filesystem read, kept as narrow as it can be
Section titled “D27 — Opening a space from the phone needs a filesystem read, kept as narrow as it can be”Every creating endpoint the app had needed something already open to hang off:
pane.split needs a pane, tab.create needs a workspace, /agent/start needs
one of those. On a Herdr session with no workspaces there was nothing to hang
off, so the phone had nothing useful to offer at all — the only way back in was
walking to the desktop. Fixing that means the app must be able to name a
directory, and it has no way to know what is on the host. Hence GET /browse.
The endpoint is deliberately the narrowest thing that can do that job:
directories only (never a file name, never contents), confined to an allowlist of
roots, and with symlinks and .. resolved and then re-checked rather than
filtered lexically — a Clean()-based check resolves .. against a symlink’s
own path instead of its target, which is a hole, and there is a test for exactly
that. A path that fails to resolve is reported as missing only when it is
nominally inside a root, so the endpoint cannot be used to probe for what exists
elsewhere on the machine.
Roots are derived, not configured: the operator’s home directory, plus the
parent of every directory Herdr already has a space open at (the point of the
flow is to open a sibling of something already open), minus anything that is an
ancestor of home — a space sitting in ~ would otherwise contribute /Users,
i.e. every account on the box.
What this is not: a defence against a compromised phone. A paired device can
already type into a shell through /send. What it avoids is a standing
disclosure of the host’s filesystem layout through a plain read endpoint — one
that needs no agent pane, is trivially harvested, and outlives revoking the
device. For the same reason POST /herdr does not clamp workspace.create’s
cwd to the roots: a device that can cd is not stopped by that, and it would
break opening a hand-typed path.
Two Herdr methods, not one, and picked by the bridge’s is_repo flag:
worktree.open for a git checkout (it attaches the repo metadata the app groups
projects by, and is idempotent — already_open instead of a duplicate space),
workspace.create for anything else (it takes any directory; worktree.open
refuses a non-repo). Verified on herdr 0.8.0: worktree.open needs cwd and
path — with path alone it answers not_git_worktree for a valid checkout.
D28 — The diff viewer derives everything it can from the payload; only the lines git never sent are an endpoint
Section titled “D28 — The diff viewer derives everything it can from the payload; only the lines git never sent are an endpoint”The Changes screen shows a directory tree, per-line numbering, word-level intra-line highlighting, and collapsed unchanged regions. Exactly one of those needed the bridge.
Everything derivable stays in the app. The tree, the hunk/line structure and
the word diff are all computed from files[].diff, which /diff already
returned (app/lib/features/diff/diff_model.dart, diff_tree.dart). Moving any
of it to the bridge would have meant a richer payload for every file of every
request — on a phone, over a tailnet — to save work the client does once per file
it actually opens, and would have coupled a rendering decision to a bridge
version. A bridge that predates all of this still renders correctly in the new
app, which is the test that matters.
The one thing that could not be derived: the unchanged lines. git diff
ships three lines of context per hunk, so everything else in a changed file is
absent from the payload. No amount of parsing recovers it. The alternative to a
new endpoint was git diff -U20 — paying for context on every file of every
request, to serve a tap most files never get, and still answering “what’s the
rest of this file?” with a bigger fixed guess. So GET /diff/expand fetches a
bounded slice per tap (see CONTRACT-diff.md), and /diff
itself is untouched.
It reads the working tree, not git history, which is only correct because of
what it is for: the endpoint fills gaps between hunks, and a line no hunk
touches is identical on both sides of the diff. That also fixes its two blind
spots by construction — a deleted file’s content exists only in HEAD, and an
untracked file’s synthetic diff already is the whole file — and the app offers
no expand affordance for either rather than one that does nothing. Same reason
an older bridge’s 404 retires the affordance silently instead of raising an
error: the honest fallback is the three lines of context git gave us, which is
what the screen showed before.
Unified, not side-by-side. At 390pt with a line-number gutter, two columns leave roughly twenty characters each — narrower than the identifiers in this repo. Diff lines wrap into a fixed gutter instead of scrolling horizontally, which also keeps the whole screen renderable as one lazy list: a horizontally scrollable code block has to lay out every line of a file to measure the widest one, which is precisely what must not happen on a several-thousand-line diff.
D29 — One mechanism answers “what can I do with this pane”
Section titled “D29 — One mechanism answers “what can I do with this pane””GET /suggestions is the single surface the app asks what a pane affords, and
internal/suggest is the single place that decides. Three features arrived
at that question separately, each with its own endpoint, its own gate and its
own affordance, and all three are now sources inside one mechanism, ranked
in one row:
| Was | Is now |
|---|---|
GET /ports + a dev-server preview chip |
the dev_server / dev_server_local sources |
GET /diff?context=1 + a “Create PR” button in the composer |
the create_pr source |
| — | the git_conflict / git_dirty / shell_idle sources |
The raw feeds survive underneath, and each still owns its own reading: /ports
is the layer that knows about lsof, HTTP probes and process trees; /diff is
the layer that knows how to run git against a pane’s cwd. The app calls neither
of them for this.
They were converging on the same question from different directions. “A server is up in this pane, here is a URL”, “this pane’s rebase stopped, here is the diff”, and “this branch has work on it, shall I open a PR” are the same sentence with different nouns. Shipping them separately would have meant three contracts, three gates, three chip-or-button surfaces, and a standing argument about which one a future affordance belongs in. The person asking does not know or care which subsystem noticed.
What made it affordable is that the scan is host-wide and cached. The
objection to merging was real: a port scan costs an lsof, a ps and a probe
per listener, and it is inherently a host question that a pane filter narrows
afterwards — so folding it into a per-pane read looks like making every pane pay
for a scan. It isn’t, because ports.Cache is shared and 5s-lived: a row of
open panes costs one lsof between them, not one each. The per-pane suggestion
cache sits just past that TTL (6s) so a miss usually finds the scan warm
rather than re-triggering one it then ignores.
One git read per pane. The same argument, made twice. Before the merge the
suggestion sources ran their own git status while the app separately polled
/diff?context=1 for the PR gate — two implementations of “what is this pane’s
git situation”, against the same directory, free to disagree about how many
files had changed. internal/gitdiff owns it now; internal/suggest receives
the answer as data and shells out for nothing. That is why the chip saying “9
files changed” and the diff screen listing nine files are the same nine.
There is one deliberate second read: the app re-checks git when a create_pr
chip is tapped. That is a pre-flight, not a gate — the bridge already decided
whether to offer it — and it exists because that action reaches outside the host
and the chip may be seconds stale.
Rank is the merge. Every source now scores itself into one ordering, and those numbers live in one block so “which of these matters more” is a single reviewable argument rather than one per feature: stuck (30) → serving (25) → changed (20) → shippable (18) → up but unreachable (15) → empty (10). A source may contribute at most two chips, so a microservice stack cannot crowd out the chip that needs a person.
kind and action stay separate, and that is what makes one mechanism
extensible. kind says why and only picks the icon; action says what and is
the only field the app branches on. An action a client does not implement — or a
known one missing its required param — is dropped rather than rendered, so a
newer bridge can add a source against an older app and the worst case is a
missing chip. It is also what makes the deferred loopback relay cheap: it
becomes one more action on an existing chip, not a second mechanism.
The mechanism carries two kinds of action, and says which
Section titled “The mechanism carries two kinds of action, and says which”Most suggestions are things the app does — open a screen, open a URL. The pull request is something only the agent can do, and folding it in without flattening that difference is the substantive part of this decision.
performer: "agent" means the app does not perform the action; it sends the
agent in the pane a message (params.prompt) asking for it. The bridge
deliberately never runs git push or gh pr create itself. A POST /pr that
shelled out to git and gh was the obvious alternative and is rejected on three
counts:
- Agent-agnostic by construction (D7). Every agent Herdr hosts has a shell. Nothing here knows or cares that it is talking to Claude.
- The agent is the one with the context. It holds the
ghauth, the repo’s commit conventions, and enough of the work to write a PR body worth reading. A bridge-side implementation would have to reinvent all three, badly. - It is watchable. Every step lands in the transcript, where it can be read and interrupted mid-flight — as opposed to an opaque HTTP call from a phone that either worked or didn’t.
Two consequences follow, both deliberate:
- The prompt is shown and editable before it is sent. Committing and pushing are irreversible and outward-facing; a phone tap that silently does them is the wrong default, and the wording is exactly what a person wants to adjust (“…and mention it supersedes #41”). The chip carries a trailing “…” for the same reason a menu item does.
- The prompt is composed on the bridge, not the client. One wording,
reviewable in one place, identical on every client — and it names the branch,
the remote and the base explicitly, because the bridge knows them and an agent
that guesses pushes to the wrong place. It is one line, because
POST /sendpastes the body and delivers Enter separately, so an embedded newline submits half a message on any agent without bracketed paste.
An absent or unrecognised performer must read as app. Failing closed is the
only safe direction: the cost of misreading an app action as an agent action is
an irreversible push nobody asked for.
What the app gave up. The PR button used to be drawn whenever the pane was in a repository at all, disabled with a sentence for the finer conditions (“This pane is on main — move the work onto a feature branch first”). As a chip it simply does not appear in those states. That is a real loss, and it is the right trade: a chip row answers what should I do now and has to stay quiet, while a sentence answers why didn’t that work — a question asked after a tap, not before one. The tap-time pre-flight keeps every one of those sentences for the case that matters most, a chip that went stale between being drawn and being tapped.
The two endpoints disagree in exactly one place, on purpose. A failed scan
is a 502 on /ports, because there the scan is the response; on
/suggestions it is logged and costs only the dev-server chips, because losing
lsof must not cost a pane its “resolve this conflict” chip.
A preview URL is built for a caller, not for the host. Found the hard way on
a real device: the dev-server URL was derived from the bridge’s own bind address,
which behind tailscale serve is 127.0.0.1:8787, so every phone was handed
http://127.0.0.1:<port> — its own loopback. The chip rendered, the browser
opened, the connection was refused. The bind address answers “where does this
process listen” and never “what should someone else dial”, and those are
different machines’ points of view.
The host now comes from the request that just succeeded (Host), falling back to
transport.public_url and then to a non-loopback bind address — first
non-loopback candidate wins, and “no candidate” means no URL rather than a link
that cannot connect. That per-caller answer is also why the suggestion cache is
keyed by client host as well as pane.
The general lesson, worth stating because the same shape recurs: this feature already reasoned carefully about bind addresses — it distinguishes a loopback-bound server from a public one and renders them differently — and then composed the public one’s URL out of the wrong machine’s address anyway. Careful reasoning about a value does not transfer to code that merely uses it. Every layer that produces a preview URL now carries a tested invariant (no loopback host may ever appear in one) rather than an intention.
A loopback-bound server is relayed, not merely explained. The bridge runs on
the Herdr host, so it can dial 127.0.0.1 when the phone cannot;
internal/preview opens a listener the phone can reach and proxies it. The chip
stops being an explanation and becomes a link.
Three choices inside that, each of which could have gone the other way:
- A listener per previewed port, not a path prefix on the bridge. A prefix
is cheaper to build and cannot be made to work: Vite, Next and Flutter web all
serve root-absolute asset paths, HMR computes its socket URL from
location, and every framework redirects to/somewhere. Making a prefix hold means rewriting HTML, CSSurl(), JS literals andLocationheaders — an arms race against every framework’s output that fails silently and differently for each. A dedicated listener gives the app a real origin, and nothing is unusual from its point of view. - The proxied request carries the target’s own
Host, not the relay’s. Dev servers increasingly refuse unknown Hosts (Vite’sallowedHosts, Django’sALLOWED_HOSTS, Rails’ host authorization), and a rejected request is a hard, confusing failure. Presenting as the local request the server already serves happily is the predictable choice; the outside authority is preserved inX-Forwarded-Host. The cost is an app that builds absolute self-URLs fromHost, which is rarer than a host check firing. - The grant is a cookie, obtained from a one-shot query parameter. The
client is a browser, not the app: it cannot set an
Authorizationheader, and it cannot attach a query parameter to the subresources the page fetches on its own. So the first navigation trades the token for anHttpOnlycookie and redirects the token out of the URL, and everything after that — including the hot-reload WebSocket handshake — carries the cookie.
What the grant is honestly worth. A paired device already has POST /send,
which is arbitrary typing into any pane, which includes curl localhost:8124.
Reaching a loopback port is not an escalation of what that bearer can do, and
the gate exists so the relay is not a hole wider than the rest of the API —
not because it is the only thing between the tailnet and this host. The relay
also binds only the address the caller reached the bridge on, so a bridge behind
tailscale serve does not put a dev server on the LAN.
Lifecycle is an idle timeout and nothing else. A dev server that dies stops being connected to, so its relay goes idle and is reaped after five minutes. Tying reaping to the port scan would couple two caches and still need this as a backstop. It also bounds the one unpleasant failure: a relay outliving its server while a different process takes that port.
A directly reachable server is never relayed. Adding a hop, a listener and a token exchange to reach something the phone can already dial is worse on every axis, so the direct URL wins whenever it exists.
What was deliberately not collapsed. /ports is not folded into
/suggestions as a ?host=1 mode: the host-wide list is a different question
with a different natural cadence, and a per-pane endpoint that sometimes answers
about the whole machine is worse than two endpoints. /diff keeps its git
object and its ?context=1 mode for the same reason — it is the diff screen’s
own header, and it is the pre-flight’s read. The scanning and the git-running
code stay in internal/ports and internal/gitdiff: internal/suggest has no
dependency beyond the standard library, and every source is a pure function of
an already-collected observation, which is what keeps adding the next one cheap.
D30 — The app speaks projects and agents; Herdr’s spaces, tabs and panes are plumbing
Section titled “D30 — The app speaks projects and agents; Herdr’s spaces, tabs and panes are plumbing”Herdr’s model is a multiplexer’s: sessions hold workspaces, workspaces hold tabs,
tabs hold panes, and a pane may or may not have a coding agent in it. That model
is correct, and the app still navigates by it — pane_id is the address /send,
/attach and /agent-transcript are keyed by, workspace_id is what
worktree.remove takes. It is simply not the model the person holding the phone
has. Theirs is: my projects, my agents, what needs me. The terminal screen was
literally titled w1N:p3.
So the ids stay on the wire and the vocabulary changes on the screen:
| Herdr | what the app says |
|---|---|
| workspace / space | project — repo + branch |
| pane running an agent | the agent |
| pane with no agent | a terminal, belonging to a project |
| tab | nothing; a grouping detail |
w1N:p3 |
never shown |
worktree.create |
start new work on a branch |
worktree.remove (+ branch delete) |
finish this work |
Three things this arrangement is deliberate about.
The mapping is a set of pure functions, in app/lib/core/naming.dart. Every
label a screen shows is derived there, so the naming can be pinned by tests
without a bridge. The way an arrangement like this fails is always the same — an
id leaks into a label — and that is invisible in review and obvious on a phone,
so test/vocabulary_test.dart and test/project_view_test.dart assert not only
what the labels say but that no pane, tab or workspace id appears anywhere on
the rendered screen.
A capability is never dropped to make the vocabulary tidy; it is re-subjected.
Tab rename is the sharp case: it renames a concept the app is hiding. It is now
offered as “Rename…” on a terminal, and only when that terminal’s tab holds
exactly one pane — which is what everything the app creates looks like — because
then renaming the tab is precisely renaming that terminal, and the subject is
one the user can see. On a split tab it is not offered, since it would silently
rename the siblings too and there is no honest label for that. tab.close
survives the same way, as “Close all in this split”. The one thing genuinely
removed is the tab.create wrapper: it backed a “New tab” menu item sitting next
to “New terminal”, which was the same action twice with Herdr’s vocabulary
leaking in to explain the difference. createPane(workspaceId:) does the same
thing and also navigates you into what it made.
Where “worktree” survives, it is because it is the precise noun. The affordances are phrased as starting and finishing work, but the branch-delete copy still names the branch, whether it is merged, and the sha a forced delete leaves behind. Those are the words that let someone get the work back, and softening them would be a worse trade than the jargon costs.
D31 — “Recently opened” is the device’s own navigation history, and cannot come from the bridge
Section titled “D31 — “Recently opened” is the device’s own navigation history, and cannot come from the bridge”recency_rank (D-era: the snapshot’s own ordering key) answers “which agent
last did something”. Home’s Recent section answers “which agent did I last
look at”. Those sound alike and are not: an agent you opened two minutes ago
belongs at the top of Recent even if it has been silent for an hour, and one
working furiously that you have never opened does not belong in it at all. No
server can know the second one — it is per device, and the bridge never sees a
navigation.
Only four fields are stored: server id, pane id, which view, when. Nothing about the agent is cached. The title, the project, the status and the age are read from live snapshot data at render, which is what stops a Recent row disagreeing with the same agent’s row three sections up — and it is also what makes the drop-out rule fall out for free: an entry with nothing live behind it resolves to nothing. A worktree removed, a pane closed, a server unpaired or merely asleep all produce the same silent shortening. There is no dead row and no error; an unreachable laptop coming back restores its rows on the next poll.
It persists in secure storage, next to the starred set (StarredAgents),
for the same reason that does: there is no schema to migrate, so it stays clear
of the drift database. The stored history is deeper (24) than the visible cap
(4), because entries drop out — a history exactly as long as the cap would show
three rows the moment one worktree was removed, with nothing to backfill from.
The recording hook lives in the screen’s State, not at the tap sites. The
honest definition of “opened” is a screen for this agent came into existence.
There are seven places that push one, so hooking the taps means a new one is a
new place to forget; hooking build unguarded would count a streaming transcript
hundreds of times a minute; hooking a snapshot refresh records the bridge
talking, not the user. A State is created once per navigation and destroyed on
pop, so a per-instance latch is exactly one visit — and re-pushing the same agent
builds a new State and records again, which is correct, because that is a
second visit. See app/lib/features/recents/record_open.dart.
D32 — Cross-origin browser clients are opt-in, and the preflight is terminated before auth
Section titled “D32 — Cross-origin browser clients are opt-in, and the preflight is terminated before auth”Hosting the web build somewhere other than the bridge — GitHub Pages, in the case that prompted this — makes the app and the API different origins, and the browser then enforces a rule the native app never met.
The failure is not obvious from either side. The app sends its bearer in an
Authorization header, which is a custom header, so before the real request the
browser sends a preflight OPTIONS. A preflight carries no credentials by
design. Every endpoint here authenticates, so the preflight was answered 401,
the browser blocked the request that would have followed, and the app — which
renders any failure as unreachable (ok => error == null) — reported the
machine as down. Meanwhile the bridge was perfectly healthy: the same URL hit
directly from the phone’s browser returned a real 401, and /snapshot with a
token returned 200 in 0.16s. “Unreachable” named the symptom and pointed at
the network, and the network was never involved.
So the preflight is terminated in middleware, ahead of the mux. It cannot be
per-handler: the whole problem is that a handler’s first act is to check a token
the preflight is not allowed to carry. A preflight is OPTIONS plus
Access-Control-Request-Method; a bare OPTIONS is an ordinary request and
still routes normally.
The allowlist is exact-match, never a wildcard. This API starts processes, reads repositories and proxies a terminal, so “which sites may script it” deserves an explicit decision.
The project’s own published UI is allowed without configuration
(config.DefaultAllowedOrigin), because it is the client this bridge exists to
serve and making every install discover and paste the same constant only turns a
working setup into a support question. Permitting an origin is not granting
access: every endpoint still demands a bearer, and CORS governs whose script
may read a reply, not who may authenticate — a visitor with no paired token can
do nothing.
transport.allowed_origins extends that set rather than replacing it, and
takes any number of entries (GOTHALO_ALLOWED_ORIGINS is the comma-separated
equivalent). Extending is what stops the obvious footgun: adding a local dev
server would otherwise silently cut off the hosted app, and the symptom would be
the same undiagnosable unreachable this whole entry is about.
Access-Control-Allow-Credentials is never sent. Auth is an explicit header,
not an ambient cookie, so there is nothing for the browser to attach on its own;
allowing credentialed requests would widen what an allowed origin can do while
buying nothing.
Note that the WebSocket endpoints (/events, /attach, the transcript stream)
were never affected: browsers do not preflight a WebSocket handshake, and
websocket.Accept already runs with InsecureSkipVerify: true, so a foreign
origin’s handshake was accepted all along. Only the HTTP surface needed this.
The app names CORS as a likely cause, hedged. A browser refuses a
disallowed response to the page, so the failure reaches Dart as status 0 with
no headers and no reason — deliberately opaque, and no probe gets around it. But
three facts are available and together they are a strong signal: we are on the
web, the bridge is a different origin from the page, and the request died with
no response at all (any reply, even a 401, proves the browser let it through).
On that the client sets BridgeException.corsBlocked and the server row reads
blocked by bridge (CORS) instead of unreachable, with the fix in the
message. The wording stays hedged — a bridge that is simply switched off looks
identical from a browser — because the point is to aim the reader at the right
machine, not to claim a certainty the platform will not give us.
