Ask ten engineers whether they have production LLM experience, and nine will say yes. Ask what broke in production, and the number drops fast. That gap — between claimed experience and verifiable experience — is the entire problem in AI hiring today, and it's the first thing our evaluation method is built to close.
Production evidence, not portfolio polish
A demo is easy to build and easy to sound enthusiastic about. A system that served real users, absorbed real traffic, and survived a real incident is not. When we screen an AI engineer, we ask about evaluation pipelines, latency and cost trade-offs, and what happened to the system months after launch — not which frameworks appear on the résumé.
- Talks about evaluation pipelines and regressions, not just prompts
- Mentions cost and latency numbers without being asked
- Has failure stories — and owns their part in them
- Knows what happened to the system months after launch
- Describes data work as most of the job, because it was
The demo tells on itself
The inverse signals are just as reliable. A candidate who name-drops frameworks with no depth on any single one, who has no real answer to “what broke?”, or who only ever reports model metrics and never business metrics has probably never operated a system under real pressure. We treat these as disqualifying, not just a yellow flag.
The third “why?” still works
Rehearsed answers survive one follow-up question. They rarely survive three. In one recent search for an LLM/NLP engineer, this single technique — asking why, and then why again — separated a shortlist of five people we'd genuinely hire from more than a hundred candidates who sounded right on paper.
What this means if you're hiring
If you're screening AI engineers yourself, the fastest improvement you can make is dropping the framework checklist and asking for the failure story instead. Anyone can list technologies. Only someone who actually ran the system can tell you what went wrong, how they found out, and what they changed because of it.