1. The argument I thought I won
A small, no-AC meeting room next to an Apple supplier line in the Yangtze Delta. 2021.
Across the table sat an Apple MQE — Manufacturing Quality Engineer, customer side. Korean. His Mandarin was better than mine.
One-on-one. No PowerPoint. Just a stack of defect photos in front of him and an open laptop in front of me.
I talked for forty minutes
I’d been doing visual QC for a little over a year. By then I’d learned the QE vocabulary cold: false positive, overkill, miss, AOQL.
I started with the convolution math. Walked through the feature maps. Explained why this specific image had been tagged NG.
Then I dropped a level — pixel-scale texture variations, gradient differences, the way a human eye actually reads a scratch.
He nodded the whole way through. Polite to a fault.
About forty minutes in, I stopped and waited.
He paused for three seconds
He looked up at me, held a three-second silence, and said — flat, calm, in a tone I still remember:
“You don’t get it — that’s why you didn’t make me get it.”
I was angry for a beat. I’ve lectured at Shanghai Jiao Tong; I know how to teach; I know every detail in this stack — what is he even talking about?
I held it together and ran the whole thing again, slower. Same result.
I walked out of that room with one private conclusion:
This MQE doesn’t get AI. He just won’t listen.
I was 36, a year-and-change out of nuclear engineering, deep in the phase where I thought I knew everything.
2. Five years later, I realized I was the one who didn’t get it
That meeting hung around for five years.
In between, I shipped a few dozen visual-defect projects. Built three systems trying to productize the same capability.
Sat through dozens more rooms — supplier lines, customer reviews, postmortems at the equipment vendor — and met more “customers who don’t get AI.” Every single time I’d file them right next to that Korean MQE.
They don’t get AI. That’s why they ask the way they do.
Late at night, his line would replay in my head. “You don’t get it — that’s why you didn’t make me get it.” It nagged. Not because it was wrong. Because I half-suspected it wasn’t.
He wasn’t challenging my technique
He wasn’t challenging my technique. He was asking a question I hadn’t noticed I was being asked:
“What gives this thing the right to judge by the same standard I do?”
He didn’t need to understand convolutions.
What he needed to understand was this: when he looks at a part with a slightly whitened edge, his head pulls up ten years of experience, this customer’s preferences, the last batch’s source material, this week’s line conditions.
What gives a network — trained on a few thousand images — the right to call NG and assume it’s the same NG?
This isn’t a technical problem
It’s whether judgment logic can be aligned.
What he wanted from me wasn’t “explain how AI works.” It was “explain how AI aligns with me.”
I didn’t say it because I couldn’t. The 2021 me didn’t yet know alignment was a thing you could engineer.
He didn’t get the technology. Fair. I didn’t get his question. Worse.
3. How we patched the gap back then — 47 if-else rules
We had an answer for that gap. It had a respectable engineering name: post-processing.
How 47 if-else rules grow
It looked like this. The model spits out a confidence score. We bolt a rule layer on top of it:
If the region’s grayscale mean is below X, area smaller than Y, and outside Z pixels from the edge → downgrade to OK.
If the texture pattern matches the one the customer flagged last week → force NG.
If the batch number starts with A8 → bump the threshold by 0.05.
...
Behind every rule sits a customer complaint, a stopped line, or a postmortem where the boss banged the table.
The 47 if-else rules I mentioned in the last post grew exactly this way.
Three chronic illnesses
The fix works — in the short run it does make the AI’s verdict look aligned with the customer. But it has three chronic illnesses:
Unmaintainable. The more rules, the tighter the coupling. A new engineer reads it once and walks away.
Non-portable. New customer, new product — the whole ruleset has to grow back from zero.
Inscrutable. You can no longer say whether the model is judging or the rules are. You’ve outsourced the alignment to 47 hand-written sentences.
We weren’t actually teaching the AI to understand the customer. We were using rules to fake understanding on the AI’s behalf.
That MQE felt this. He couldn’t put it in those words. But he knew — what you’re aligning isn’t me.
4. Five years on, the patch has a new name: “the LLM”
In 2026 the language is unrecognizable.
We don’t say “post-processing” anymore. We say priors.
We don’t say “add a rule.” We say put it in the system prompt.
We don’t say “bolt judgment logic on top.” We say use a VLM for semantic understanding.
New language, same plumbing
What’s the underlying activity? The same thing as 2021 — patching what the AI hasn’t seen in its training data.
Just that the patch material has changed. From hardcoded if-else to a model pre-trained to “know a little about everything.”
The advantages are real:
It “understands” customer intent instead of just matching thresholds.
It carries common sense across products without per-line rule rewrites.
It can explain itself — at minimum, in something that reads like a human sentence.
But the substance hasn’t moved.
It’s still a patch — supplying what the AI’s training data couldn’t supply on its own: a human judgment standard.
2021 vs. 2026
The only real differences:
When Patch material Patch boundary Failure mode 2021 Post-processing rules This customer’s recent gripes One ruleset per customer, growing tangled 2026 Priors / LLMs / world knowledge Internet-scale common sense Looks universal; still needs per-customer calibration
The progress isn’t “we have LLMs now so the patch is gone.” The progress is that we finally understand the patch should be a designed part of the system, not a duct-tape strip slapped on after the fact.
This thing has been named, productized, argued about.
But there are still plenty of projects out there patching the 2021 way. Same melody, bigger model — the if-else just slid into the system prompt. The ending can be identical:
47 if-else rules become 47 system prompts. Same outcome.
5. What that MQE was actually asking — said properly
Look back at that 2021 afternoon.
The man had no AI training. His toolkit was sampling plans, SPC charts, 5-Whys, fishbone diagrams. How did he see through me in three seconds?
How he made the read
Because aligning human judgment is what he’d done his whole working life.
The entire point of a quality engineer is to take “OK / NG” — a subjective call — and turn it into something repeatable, auditable, and transferable across different eyes, different batches, different customers.
He listened to me for forty minutes and what he heard wasn’t “how AI works.” What he heard was:
“This thing can’t carry the judgment system I’ve spent ten years building.”
Three minimum bars for a stable judgment system
What he saw through wasn’t whether I understood AI. It was whether I understood what a stable judgment system looks like.
The minimum bars:
Same part, today and tomorrow → same call.
Different lines, same standard → comparable results.
Something goes wrong → you can trace it to a specific decision criterion, not “the model said so.”
My 2021 version failed all three. The 47-rule version failed all three.
And today — if you’re not careful — the 47-system-prompt version still fails all three.
What that MQE refused on instinct is the exact problem the entire industrial AI industry still hasn’t solved.
He wasn’t wrong. I was. I thought he was challenging my explanation. He was challenging my approach.
6. What I’d say to that Apple MQE today
If I could sit back down in that hot little meeting room, I wouldn’t talk about convolutions. I wouldn’t talk about transformers.
I’d say:
You said I didn’t get it. You were right.
Back then I thought I did — but really, I’d just learned to dress up “humans-in-the-loop” as “AI making the call.” I couldn’t answer your real question: when I hand this thing over to you, can it judge the way you judge, write down its judgments the way your team does, pass them on to your successor the way your quality system does?
Our industry has spent the last five years slowly learning to answer that one question. We’ve gone from rules, to priors, to LLMs, to multimodal alignment — every step making “aligning human judgment” a little more systematic, a little more designable.
I still can’t put my hand on my heart and say I’ve solved it. What I have learned is that the question itself deserves to be taken seriously — and the 2021 version of me didn’t even see it as a question.
Thank you for that one sentence. It took me five years to start asking the right one.
7. Five questions to check yourself against
If you’re working on an industrial AI project in any role — customer, vendor, engineer, decision-maker — run these five:
Once your model gives a verdict, can you explain how it’s aligned to the customer’s standard?
When the customer reports a new defect type, is your first instinct to fix the data, fix the process — or to add a rule?
How many human-written backstops are in your system right now? Can anyone still describe how they interact?
If a new customer QE took over tomorrow, could they independently audit whether your system is fit for purpose within 90 days?
The “LLM-powered” line on your marketing page — is the system actually understanding the customer, or have you just rewritten the rules in a different language?
Three or more of these unanswered — you’re probably walking the same path 2021 me walked.
8. Next post
The next piece digs a layer deeper: “priors / world knowledge / LLMs” — those patches have to live somewhere. Which layer of the system? Maintained by whom? How do you keep them from becoming the new 47 system prompts?
The answer isn’t in the model. It’s in the architecture — there are three architecture choices industrial software should have made ten years ago. Get them wrong and AI stays stapled to the UI. Get them right and the day an agent takes over your system, you don’t have to rewrite anything.
If you’re working on this,
— let’s walk this road together.
— Yuanwei, from a robotics lab in Novi, Michigan
April 28, 2026





