Content hub Back to the console

Engineering posts: six publish-ready drafts

General wisdom from real work. No system names, no topology, no operational status. Each post stands alone. Rotate one per account per day, adapt the opener per network.


E1 · Silence is not health

Tags: programming, devops, monitoring

Every ops team thinks they want a quiet dashboard. They don't. They want an honest one, which is a different thing.

We had a status file that updated every cycle. One night the process writing it died. The file kept serving its last known state, and for three and a half hours everything downstream reported green, because green was simply the last thing written.

The fix wasn't clever. A heartbeat with an expiry: if the file is older than two cycles, it reads as stale, and stale renders as broken, never as healthy. A lock so two writers can't race each other. And a log line at the start of every cycle, not just the end, because a log that only writes on success tells you nothing about the failure you're currently inside.

The general rule we kept: a monitor that dies silently is worse than no monitor at all. No monitor leaves you suspicious. A dead monitor leaves you confident.

If your system has a health endpoint, ask it one question today. What happens when the thing that writes the health data stops running? If the answer is "it keeps serving the last value," you don't have a health check. You have a timer with good PR.


E2 · Idempotency is a money feature

Tags: programming, blockchain, engineering

Here is a bug pattern that costs real money, and it almost never shows up in tutorials.

An operation fires. The network is slow. The caller sees no response, assumes failure, and retries. Except the first call succeeded. Now the operation happened twice. If that operation moves value, someone just paid double, and your logs will happily tell you both calls "worked."

The root cause is almost always the same: a fresh identifier generated per attempt. Request IDs built from timestamps or random values guarantee that the retry looks brand new to any deduplication ledger downstream. The ledger can't recognize a repeat it was never given a stable key for.

The fix is boring and it works: derive the request identifier from the operation's content, not from the attempt. Same payload, same ID, every time. Now the retry hits the ledger, the ledger says "already processed," and the caller gets the original result instead of a duplicate effect.

Second half of the fix: classify rejections structurally. A 400, a 403, an explicit chain-level rejection means never auto-retry. Retrying a rejection is how you turn one honest failure into twenty confused ones.

If your system retries anything that touches money, go read your request ID generation right now. If it contains Date.now() or Math.random(), you have a double-spend waiting for a slow network day.


E3 · Our cost model was wrong by 3.5x, and the fix was embarrassingly simple

Tags: blockchain, tron, measurement

We needed to know what a small memo costs on a DPoS chain with a bandwidth economy. So we built a model. Careful one. Accounted for message size, energy versus bandwidth pools, the whole structure. The model said: roughly 0.6 units per memo.

Then we broadcast one and read the receipt. 2.1 units. Not 20% off. Three and a half times off.

The instinct was to debug the model. Add terms, tune coefficients. We didn't. We did the dumb thing first: we collected about a dozen real receipts from the chain, spread across hours, and plotted actual cost against payload size. The receipts didn't lie and they didn't fluctuate much. Roughly eight bandwidth points per byte, plus a floor that our model hadn't priced because the floor isn't in the documentation where you'd expect it.

The lesson isn't about that chain. It's about the order of operations. Model first, measure never, is how confident people publish numbers that are wrong by 250%. Measure first, model second, and the model becomes a compression of reality instead of a guess wearing a lab coat.

We published the corrected number along with the old one and the gap between them. Weirdly, that's when people started trusting our other numbers.


E4 · Fail closed, or fail honest. There is no third option.

Tags: engineering, security, design

Systems fail. That part is not optional. What is optional is the style of the failure, and almost every bad outage I've studied came from a system that failed in a third style: it guessed.

Fail closed means: when the input is wrong, missing, or ambiguous, the operation refuses and says why. The key doesn't verify against the chain? No broadcast. The balance can't be confirmed? No transaction. The refusal is the feature. Nothing moves until the world makes sense.

Fail honest means: when the operation can't complete, the system reports exactly that, with the reason, instead of reporting success and hoping. A status endpoint that returns a cached green over a dead dependency isn't lying with intent, but the effect on whoever trusts it is identical.

Guessing is the third style and it's the killer. The library returns undefined, so the code picks a default. The endpoint times out, so the retry assumes the previous state. The validator can't find the record, so it treats absence as approval. Every guess is a small bet that nobody wrote down, placed with money that isn't the guesser's.

Our internal rule after one too many postmortems: every ambiguous branch must resolve to either a refusal or an explicit "unknown" that propagates upward. Defaults are banned at boundaries. It makes the system feel stricter, because it is. The trade is worth it. A strict system annoys you on Tuesday. A guessing system bills you on Friday.


E5 · The bug was in the test

Tags: programming, testing, python

We wrote a test suite for a signing helper. Twenty-odd assertions, all green in development. In production the helper failed immediately, and the suite still said green.

The reason took an afternoon to find and two minutes to state: the tests were written against the API we remembered, not the API that shipped. Function names had drifted. A parameter that the real code required had a default in our test's imagination. The suite wasn't testing the module. It was testing a fan-fiction version of the module that existed only in the test file.

The fix wasn't rewriting assertions. It was adding a calibration step: before any behavioral test runs, introspect the actual module. Confirm the functions exist. Confirm the signatures. Confirm the return shapes against real calls with real fixtures. Only then run the behavior tests.

We now treat "the test passed" as a claim that needs its own evidence. Every suite starts by proving it is pointed at reality. It sounds paranoid right up until the first time it catches a rename that would have shipped.

If you maintain tests for anything important, try this today: grep your test files for function names, then grep the source for the same names. Any name that appears only in the tests is a ghost. Ghosts don't fail loudly. They pass.


E6 · Frozen text is a lie waiting to happen

Tags: programming, ui, honesty

Somewhere in your interface there is a string that was true when someone typed it and has been decaying ever since. "Created in 2018." "Over 1,000 users." "Backed by X." Nobody checks it. It renders on every load, in every state, forever.

We found ours during an audit: a claim about account origins, hardcoded into a component three iterations ago. The data field it described had gained new entries since. The string kept insisting on the old story because strings never update themselves.

The rule we adopted has a name now: derived or silence. Any claim in the UI is either computed from live data at render time, or it doesn't appear. The "created" line now derives from the account records, and the moment a record contradicts it, the line changes or disappears. No human has to remember to update it, because there is nothing left to update.

The second half of the rule matters as much: when live data isn't available, the honest state is "unknown," rendered as such. A skeleton, a dash-free placeholder, a "pending first measurement" note. Anything except a confident sentence backed by nothing.

Go find your frozen strings today. Search the codebase for years, for round numbers, for the word "over." Each hit is a small timer counting down to the day a user notices your interface lying to them, and believes every other number a little less.