Field Service
The lab is where things work. The field is where you find out which of your assumptions were load-bearing.







Some of it doesn't happen locally. Stendal, in Germany — a machine, a hotel, and a fortnight to make it run. The work is identical and everything around it isn't: someone else's voltage, someone else's spares, a supplier three time zones away, and no possibility of driving back for the part you forgot.
It sharpens you. You stop assuming anything can be fetched.




Every one of these was taken so I could look at it later without the machine in front of me. Photographing what you're about to take apart is the cheapest insurance in the trade, and it doubles as the report you'll be asked for a week after you've forgotten.
In fairness to the job: the machine is only there five days a week, and you are there for two or three weeks. Nobody puts this part in the quote either.







Two weeks away from your family, someone else’s voltage, and a 3am phone call you can’t take because of the time difference. And also: Berlin on a Sunday, and a table of people who were strangers on Monday. Both halves are true and neither cancels the other.
Not every call is electrical. Sometimes the answer is a worn bore, a seat that isn't sealing, or a shaft that has been running out of true for long enough to leave the evidence on itself.
This is a different mode of work — slower, closer, and mostly photography. You are building a case, not fixing a fault, and the case has to hold up in a room where somebody would prefer it were nobody's fault.




Nearly everyone debugs by pattern-matching to the last thing that broke. It works often enough to be a trap. The discipline is refusing to touch anything until you can state what the machine is actually doing, in terms that could be wrong.
- 1Reproduce it before you theorize. A fault you can't summon on demand is a rumor. Half of all field calls end here, because the act of reproducing it reveals the cause.
- 2Read the machine, not the complaint. "It just stops randomly" has never once been random. It's a sensor, a timeout, or a thing that happens every 40 minutes that nobody connected to the failure.
- 3Change one thing. Change two and you have learned nothing, whichever way it goes. This is the rule everyone knows and nobody follows at hour eleven.
- 4The problem is almost never the PLC. It's a field device, a connection, a mechanical fault, or an assumption. I've been doing this since 1999 and the controller was genuinely at fault a number of times I could count without taking my boots off.
- 5Write down what you found. Not for them. For the next person, who is statistically likely to be you.
Staying methodical while someone senior stands behind you asking how long. That's the job. The technical part is comparatively easy and I've never once been paid for the technical part alone.
A panorama sweep that assumed the world would hold still. It did not. The algorithm made exactly one bad assumption and then committed to it, with total confidence, all the way across the frame.
Which is the whole argument for stating a tolerance: every measurement carries one, whether or not anybody wrote it down.
SCALE: NONENOT TO BE USED FOR FABRICATION