All posts
·12 min read

The Environment You Could Only Run by Breaking the App

Local development was maintained as a ten-patch diff against product source, three of which disabled authentication. The patches were not the problem. They were the symptom of an environment that could not be bootstrapped.

  • #PlatformEngineering
  • #DeveloperExperience
  • #Infrastructure
  • #Testing

Most codebases that have been worked on for a couple of years grow a document with a name like DEV-HACKS.md. Ours had ten entries. Three of them disabled authentication.

They were not written by careless people. Each one had been discovered by somebody hitting a wall at nine at night, working out the minimum edit that got past it, and being decent enough to write it down for whoever hit the same wall next. The file was an act of generosity. It was also a set of instructions for modifying product source in order to run product source, which meant every engineer's working tree contained changes that must never be committed, and the first question on any strange bug was "is that real, or is that your patches?"

I spent August replacing that file. What I want to write about is not the tool that replaced it, but the one constraint that decided everything about the tool's shape.

Why the patches existed

The obvious explanation is drift: nobody kept the local setup working, so people improvised. That explanation is comfortable and it was wrong.

The environment could not be bootstrapped at all. The schema was managed by migrations, and not one of the one hundred and sixty-three migrations contained a single INSERT. Migrations built tables. Nothing built rows. Separately, the local identity provider came up with its users and groups but no role bindings, so even once you had an account there was nothing it was permitted to do.

So a clean clone gave you a correct, empty, unusable system. There was no sequence of documented steps that produced a working application, which meant the only remaining move was to edit the application until it stopped asking for the things that were missing. Three patches disabling auth is not what happens when developers are lazy. It is what happens when authentication is the first thing that fails and nobody can seed an authorisation model.

Reverting all ten patches was the first commit of the project, and it was the useful one. It did not fix anything. It made every real dependency fail loudly and in order, which is the only honest starting inventory you can get.

The rule that shaped everything

There was one hard constraint, and it was mine rather than anybody else's:

The tool may not change the application.

Not "should avoid". May not. It owns command strings, environment variables and fixtures. It does not own a single line of product source, and if the only way to make something work is a patch, then that thing does not work.

This sounds like an ideological position and it is really a practical one. A tool that is allowed to edit the app is a patch set with better ergonomics. It rots the moment the app moves, its failures are indistinguishable from product bugs, and it cannot be handed to somebody else, because handing it over means handing over a fork. Forbidding the edit is what turns a script into something a second person can rely on.

Nearly every design decision below is downstream of that one rule.

Redirecting an application you are not allowed to edit

If you cannot change the app, the only lever left is what it reads from its environment. So components declare what they are, and the tool derives what the application needs to be told:

components:
  mysql:
    kind: container
    port: 3306
    provides: [database]
  casdoor:
    kind: container
    port: 8000
    provides: [identity]
  signoz:
    kind: compose
    port: 4318
    provides: [observability]

A component that provides: [database] causes the resolver to emit the database host and port. One that provides identity emits the issuer base URL. Nothing in the manifest names a variable, because the mapping from capability to variable name is the tool's job, and it is the part that has to stay correct as components move around.

Which brings me to the bug I liked least.

The derivation emitted DB_HOST and DB_PORT. Sensible names. Nothing in the application reads them. It reads MYSQL_HOST and MYSQL_PORT. The result was that you could place the database on another machine, watch it start there, watch the seed run there against the correct tables, and the backend would quietly go on talking to localhost for the entire session. Placement appeared to work perfectly and did nothing whatsoever.

I have written before about a delivery layer that reported total success while delivering nothing, and this is the same mistake wearing different clothes. Both systems reported on their own intentions. Neither checked the only thing that mattered, which was what happened at the other end. The fix was one line. Proving the fix meant opening a connection with the application's real credentials and confirming which of two containers answered.

The general rule I took from it: grep for the variable the application actually reads before you name one you are going to derive. There is no error for supplying configuration nobody consumes.

There is a related subtlety worth knowing if you try this. The tool sets thirteen variables. The committed .env.local file still governs the other hundred and five, and precedence works only because dotenv.config() does not overwrite a value already present in process.env. That is load-bearing behaviour in a third-party library, documented but easy to miss, and the entire override model rests on it.

What is shared, and what is personal

The second consequence of the rule is a split that I think is the most portable idea in the whole thing.

devplane.yaml is committed and identical for everyone. It describes what the environment is: the components, their ports, their capabilities, their seed stages. It contains no absolute paths and no machine addresses, because those are not facts about the project.

~/.devplane/profile.yaml is per-engineer and never committed. It describes where things run:

name: mithilesh

hosts:
    remote-box:
        address: 10.0.0.107
        user: root

placement:
    signoz: { location: remote, host: remote-box }

Machines are declared once and referenced by name, so an address appears in exactly one place. Anything not named stays local. And for the case that is not a standing decision, a flag that outranks the profile for one command only:

devplane start --at mysql=remote-box
devplane start --at signoz=local        # the reserved undo

The environment definition is shared. Placement is personal. Once that line is drawn, a stack of arguments stop happening. Whether the observability stack should run on your laptop is no longer a team decision requiring consensus, because it was never a property of the project. Mine runs on a remote box permanently and takes three and a half gigabytes of memory with it. Somebody with more RAM and no spare machine runs it locally. Neither of us edits a shared file, and neither of us is wrong.

Remote components are driven over an SSH Docker transport, which means one code path serves both placements. Where an open port is undesirable, a component can be tunnelled instead: the SSH forward is opened before the health check and closed on stop, and the application is still told localhost, because from its point of view that is true.

Knowing what must not move

A tool that can relocate things needs to know what it may not relocate, and it needs to be able to say why.

Four components are pinned. The backend, the serverless emulator and the crypto service need VPN routing and host file entries; the frontend dev servers are browser-facing. The resolver refuses to move them, and it refuses with the reason attached rather than with a generic error.

There is a fifth state that I have come to think is underrated. A component can be declared handoff, which means the tool will not start it, will reserve its port so nothing else takes it, and will print the command you should run yourself. That is the mode for the thing you are actively debugging in an IDE, attached to a profiler, or running from a branch the tool knows nothing about.

Platforms are usually judged on what they automate. A good one also has a sanctioned way to get out of the way, because otherwise the first engineer with an unusual need goes around it, and once one person is outside the tool the tool is no longer describing reality.

Testing something that starts containers, without starting containers

The core of the tool is one function:

(manifest, profile, flags) -> plan

It is pure. It reads no disk beyond the two config files, touches no network, starts nothing. Everything interesting lives there: placement, pin enforcement, capability derivation, port clash detection across machines, mode selection. Executing the plan is a thin layer of adapters underneath.

The payoff is that every combination of placement and mode is a unit test that runs in milliseconds with no container runtime present. Seventy-seven of them run on a laptop with Docker stopped. If the resolver had been fused to execution, the only way to test "what happens when the database is remote and tunnelled and the backend is pinned local" would have been to actually do it, which is slow enough that in practice nobody would have.

I did not appreciate at the time that this decision is also what makes the tool explicable. Because a plan is a value, it can be printed, and "show me what you are about to do" turns out to be the thing people ask for first.

Seeding is the actual product

The orchestration was the interesting part to build. The seeding is the part that made the difference, and it is where the original problem actually lived.

Seven stages, each independently runnable, each idempotent, so that a second run is a no-op rather than a mess:

  • the database baseline, restored only when the database is empty
  • reference data, the rows the migrations never carried
  • the identity dump, verified against a recorded hash so a drifted copy fails loudly instead of producing a subtly different permission set
  • object storage, mirrored byte-identical
  • encryption secrets
  • parameter store entries
  • observability registration, which is load-bearing and looks like an afterthought

The one I would defend hardest is the secrets stage. Thirty-seven per-tenant encryption secrets are needed for the crypto service to resolve anything. None of them are copies. Each is derived deterministically from the tenant name and a fixed local prefix, so every engineer's environment agrees with every other engineer's, the values are stable across rebuilds, and no real secret has ever left the boundary it belongs in. Local development is one of the great quiet sources of credential sprawl, and "derive something consistent" is nearly always available where "copy the real one" is the default.

The idempotency is not fastidiousness either. It is what makes the rest of this post's ending possible.

What it still does not do

Two honest gaps.

Data does not follow placement. Volumes belong to the machine they were created on, and stopping a component never removes them. Moving the database to another machine switches datasets; it does not migrate anything, and moving back returns the original data untouched. This is defensible and occasionally surprising, and there is no migrate command.

Nothing is provisioned. There is no Terraform, no cloud-init, no image build. The tool assumes a machine that already exists, reachable, with a Docker engine and passphrase-less SSH key auth. The setup guide says so plainly and then gives you the manual steps, which is the honest version of admitting that the first fifteen minutes of the story are missing.

Terraform is the obvious next step, and the interesting question is not whether to use it but where to stop.

The temptation is to go all the way. There is a Docker provider, it speaks to a remote engine over SSH, and you could declare the containers themselves as resources. I do not think that is right, because it would mean modelling start as apply and stop as destroy, and destroy takes the volumes with it. The guarantee that stopping your environment never costs you your data is the single property I would least like to trade. Terraform also has no vocabulary for a component that should be started by a human in an IDE, or for hot reload, and provisioners are a last resort by its own documentation, so the seed pipeline is staying where it is.

The line I would draw is this. Terraform owns what has drift. The tool owns what has a session. The machine, its firewall rules, its key material, its Docker engine: those are infrastructure, they persist, they can rot silently, and they want reconciliation. Two output values, an address and a user, and they land in the hosts block of a personal profile. Everything above that line is a working day, and a working day is not a desired state.

What makes this pleasant rather than daunting is the seeding work, which is why it was worth doing properly. Because all seven stages are idempotent, a provisioned box is disposable in a way a hand-built one never is. Destroying it costs the minutes it takes to seed a new one, and nothing else. The unglamorous fixture pipeline is what earns the right to treat the machine as cattle.

The thing I should have done first

I never measured the baseline.

I know the patch count went from ten to zero, because that is countable after the fact. I do not know what a clean clone to a first authenticated request cost before, in minutes or in manual steps, because by the time it occurred to me to measure it the old environment no longer existed to measure. Every claim I can now make about the improvement is a claim about the shape of the work rather than its size.

If you are about to replace something painful, spend the first hour timing the pain. It is the cheapest hour in the project and it is the only one you cannot go back for.

The shape of the lesson

The tool is a manifest, a resolver, four adapters and a seed pipeline, and none of that is novel. The decision that made it work was refusing to let it touch the thing it manages, and then discovering that every subsequent question had an answer implied by that refusal.

A platform is not defined by what it automates. It is defined by what it is forbidden to touch, because that boundary is the only reason anyone else can trust it.