The Environment You Could Only Run by Breaking the App
Local development was maintained as a ten-patch diff against product source, three of which disabled authentication. The patches were not the problem. They were the symptom of an environment that could not be bootstrapped.
- #PlatformEngineering
- #DeveloperExperience
- #Infrastructure
- #Testing
Most codebases that have been worked on for a couple of years grow a document
with a name like DEV-HACKS.md. Ours had ten entries. Three of them disabled
authentication.
They were not written by careless people. Each one had been discovered by somebody hitting a wall at nine at night, working out the minimum edit that got past it, and being decent enough to write it down for whoever hit the same wall next. The file was an act of generosity. It was also a set of instructions for modifying product source in order to run product source, which meant every engineer's working tree contained changes that must never be committed, and the first question on any strange bug was "is that real, or is that your patches?"
I spent August replacing that file. What I want to write about is not the tool that replaced it, but the one constraint that decided everything about the tool's shape.
Why the patches existed
The obvious explanation is drift: nobody kept the local setup working, so people improvised. That explanation is comfortable and it was wrong.
The environment could not be bootstrapped at all. The schema was managed by
migrations, and not one of the one hundred and sixty-three migrations contained
a single INSERT. Migrations built tables. Nothing built rows. Separately, the
local identity provider came up with its users and groups but no role bindings, so
even once you had an account there was nothing it was permitted to do.
So a clean clone gave you a correct, empty, unusable system. There was no sequence of documented steps that produced a working application, which meant the only remaining move was to edit the application until it stopped asking for the things that were missing. Three patches disabling auth is not what happens when developers are lazy. It is what happens when authentication is the first thing that fails and nobody can seed an authorisation model.
Reverting all ten patches was the first commit of the project, and it was the useful one. It did not fix anything. It made every real dependency fail loudly and in order, which is the only honest starting inventory you can get.
The rule that shaped everything
There was one hard constraint, and it was mine rather than anybody else's:
The tool may not change the application.
Not "should avoid". May not. It owns command strings, environment variables and fixtures. It does not own a single line of product source, and if the only way to make something work is a patch, then that thing does not work.
This sounds like an ideological position and it is really a practical one. A tool that is allowed to edit the app is a patch set with better ergonomics. It rots the moment the app moves, its failures are indistinguishable from product bugs, and it cannot be handed to somebody else, because handing it over means handing over a fork. Forbidding the edit is what turns a script into something a second person can rely on.
Nearly every design decision below is downstream of that one rule.
Redirecting an application you are not allowed to edit
If you cannot change the app, the only lever left is what it reads from its environment. So components declare what they are, and the tool derives what the application needs to be told:
components:
mysql:
kind: container
port: 3306
provides: [database]
casdoor:
kind: container
port: 8000
provides: [identity]
signoz:
kind: compose
port: 4318
provides: [observability]
A component that provides: [database] causes the resolver to emit the database
host and port. One that provides identity emits the issuer base URL. Nothing in
the manifest names a variable, because the mapping from capability to variable
name is the tool's job, and it is the part that has to stay correct as components
move around.
Which brings me to the bug I liked least.
The derivation emitted DB_HOST and DB_PORT. Sensible names. Nothing in the
application reads them. It reads MYSQL_HOST and MYSQL_PORT. The result was
that you could place the database on another machine, watch it start there, watch
the seed run there against the correct tables, and the backend would quietly go on
talking to localhost for the entire session. Placement appeared to work
perfectly and did nothing whatsoever.
I have written before about a delivery layer that reported total success while delivering nothing, and this is the same mistake wearing different clothes. Both systems reported on their own intentions. Neither checked the only thing that mattered, which was what happened at the other end. The fix was one line. Proving the fix meant opening a connection with the application's real credentials and confirming which of two containers answered.
The general rule I took from it: grep for the variable the application actually reads before you name one you are going to derive. There is no error for supplying configuration nobody consumes.
There is a related subtlety worth knowing if you try this. The tool sets thirteen
variables. The committed .env.local file still governs the other hundred and
five, and precedence works only because dotenv.config() does not overwrite a
value already present in process.env. That is load-bearing behaviour in a
third-party library, documented but easy to miss, and the entire override model
rests on it.
What is shared, and what is personal
The second consequence of the rule is a split that I think is the most portable idea in the whole thing.
devplane.yaml is committed and identical for everyone. It describes what the
environment is: the components, their ports, their capabilities, their seed
stages. It contains no absolute paths and no machine addresses, because those are
not facts about the project.
~/.devplane/profile.yaml is per-engineer and never committed. It describes
where things run:
name: mithilesh
hosts:
remote-box:
address: 10.0.0.107
user: root
placement:
signoz: { location: remote, host: remote-box }
Machines are declared once and referenced by name, so an address appears in exactly one place. Anything not named stays local. And for the case that is not a standing decision, a flag that outranks the profile for one command only:
devplane start --at mysql=remote-box
devplane start --at signoz=local # the reserved undo
The environment definition is shared. Placement is personal. Once that line is drawn, a stack of arguments stop happening. Whether the observability stack should run on your laptop is no longer a team decision requiring consensus, because it was never a property of the project. Mine runs on a remote box permanently and takes three and a half gigabytes of memory with it. Somebody with more RAM and no spare machine runs it locally. Neither of us edits a shared file, and neither of us is wrong.
Remote components are driven over an SSH Docker transport, which means one code
path serves both placements. Where an open port is undesirable, a component can be
tunnelled instead: the SSH forward is opened before the health check and closed on
stop, and the application is still told localhost, because from its point of view
that is true.
Knowing what must not move
A tool that can relocate things needs to know what it may not relocate, and it needs to be able to say why.
Four components are pinned. The backend, the serverless emulator and the crypto service need VPN routing and host file entries; the frontend dev servers are browser-facing. The resolver refuses to move them, and it refuses with the reason attached rather than with a generic error.
There is a fifth state that I have come to think is underrated. A component can be
declared handoff, which means the tool will not start it, will reserve its port
so nothing else takes it, and will print the command you should run yourself. That
is the mode for the thing you are actively debugging in an IDE, attached to a
profiler, or running from a branch the tool knows nothing about.
Platforms are usually judged on what they automate. A good one also has a sanctioned way to get out of the way, because otherwise the first engineer with an unusual need goes around it, and once one person is outside the tool the tool is no longer describing reality.
Testing something that starts containers, without starting containers
The core of the tool is one function:
(manifest, profile, flags) -> plan
It is pure. It reads no disk beyond the two config files, touches no network, starts nothing. Everything interesting lives there: placement, pin enforcement, capability derivation, port clash detection across machines, mode selection. Executing the plan is a thin layer of adapters underneath.
The payoff is that every combination of placement and mode is a unit test that runs in milliseconds with no container runtime present. Seventy-seven of them run on a laptop with Docker stopped. If the resolver had been fused to execution, the only way to test "what happens when the database is remote and tunnelled and the backend is pinned local" would have been to actually do it, which is slow enough that in practice nobody would have.
I did not appreciate at the time that this decision is also what makes the tool explicable. Because a plan is a value, it can be printed, and "show me what you are about to do" turns out to be the thing people ask for first.
Seeding is the actual product
The orchestration was the interesting part to build. The seeding is the part that made the difference, and it is where the original problem actually lived.
Seven stages, each independently runnable, each idempotent, so that a second run is a no-op rather than a mess:
- the database baseline, restored only when the database is empty
- reference data, the rows the migrations never carried
- the identity dump, verified against a recorded hash so a drifted copy fails loudly instead of producing a subtly different permission set
- object storage, mirrored byte-identical
- encryption secrets
- parameter store entries
- observability registration, which is load-bearing and looks like an afterthought
The one I would defend hardest is the secrets stage. Thirty-seven per-tenant encryption secrets are needed for the crypto service to resolve anything. None of them are copies. Each is derived deterministically from the tenant name and a fixed local prefix, so every engineer's environment agrees with every other engineer's, the values are stable across rebuilds, and no real secret has ever left the boundary it belongs in. Local development is one of the great quiet sources of credential sprawl, and "derive something consistent" is nearly always available where "copy the real one" is the default.
The idempotency is not fastidiousness either. It is what makes the rest of this post's ending possible.
What it still does not do
Two honest gaps.
Data does not follow placement. Volumes belong to the machine they were created on, and stopping a component never removes them. Moving the database to another machine switches datasets; it does not migrate anything, and moving back returns the original data untouched. This is defensible and occasionally surprising, and there is no migrate command.
Nothing is provisioned. There is no Terraform, no cloud-init, no image build. The tool assumes a machine that already exists, reachable, with a Docker engine and passphrase-less SSH key auth. The setup guide says so plainly and then gives you the manual steps, which is the honest version of admitting that the first fifteen minutes of the story are missing.
Terraform is the obvious next step, and the interesting question is not whether to use it but where to stop.
The temptation is to go all the way. There is a Docker provider, it speaks to a
remote engine over SSH, and you could declare the containers themselves as
resources. I do not think that is right, because it would mean modelling start
as apply and stop as destroy, and destroy takes the volumes with it. The
guarantee that stopping your environment never costs you your data is the single
property I would least like to trade. Terraform also has no vocabulary for a
component that should be started by a human in an IDE, or for hot reload, and
provisioners are a last resort by its own documentation, so the seed pipeline is
staying where it is.
The line I would draw is this. Terraform owns what has drift. The tool owns what
has a session. The machine, its firewall rules, its key material, its Docker
engine: those are infrastructure, they persist, they can rot silently, and they
want reconciliation. Two output values, an address and a user, and they land in
the hosts block of a personal profile. Everything above that line is a working
day, and a working day is not a desired state.
What makes this pleasant rather than daunting is the seeding work, which is why it was worth doing properly. Because all seven stages are idempotent, a provisioned box is disposable in a way a hand-built one never is. Destroying it costs the minutes it takes to seed a new one, and nothing else. The unglamorous fixture pipeline is what earns the right to treat the machine as cattle.
The thing I should have done first
I never measured the baseline.
I know the patch count went from ten to zero, because that is countable after the fact. I do not know what a clean clone to a first authenticated request cost before, in minutes or in manual steps, because by the time it occurred to me to measure it the old environment no longer existed to measure. Every claim I can now make about the improvement is a claim about the shape of the work rather than its size.
If you are about to replace something painful, spend the first hour timing the pain. It is the cheapest hour in the project and it is the only one you cannot go back for.
The shape of the lesson
The tool is a manifest, a resolver, four adapters and a seed pipeline, and none of that is novel. The decision that made it work was refusing to let it touch the thing it manages, and then discovering that every subsequent question had an answer implied by that refusal.
A platform is not defined by what it automates. It is defined by what it is forbidden to touch, because that boundary is the only reason anyone else can trust it.