Building AI software got cheap and proving it works became the jobThe development team has become the QA team. It has five things to prove.The Applied AI Summit, which began in 2020 as the NLP Summit, has always been the conference for practitioners: the state of the art in industry, as a complement to the conferences that present the academic state of the art. For six years, people presented variants of how to build useful things: fine-tuning deep learning models, optimizing RAG pipelines, scaling real-world systems at a reasonable cost. This year I was surprised to find that about two thirds of the submissions were about how to validate and test AI systems. They came from every industry and from companies of every size. Building now takes weeks but validation still takes quartersDeveloping software has become far easier and faster. You can now build complex software without experts in the minutiae of front-end libraries, DevOps tools, or back-end scaling and partitioning, which is what developers used to specialize in. One engineer can build an end-to-end, scalable, industry-specific customer service platform, HR platform, or medical decision support system in a few weeks. One of this year’s keynotes, on the infrastructure behind production agents from ThoughtSpot, describes the same speed for AI agents: adding one takes a configuration entry and weeks of effort, where it used to take a codebase and quarters. What remains hard is validating what was built: someone still has to write the test cases, a domain expert still has to say what a correct answer is, and a person still has to read the failures. A 2026 economics working paper describes the two curves: the cost to automate a task keeps falling, and the cost to verify the output is limited by how much people can check. A summit session on comprehension debt describes the same gap inside engineering teams: AI generating code, prompts, and integrations faster than engineers can understand or validate them. Letting the AI check itself does not close the gap. A BlackRock session on governing AI-assisted software development covers AI writing both the code and its tests from the same flawed assumptions, so defects pass validation. Someone still has to show that the system won’t make a catastrophic financial or legal mistake. An insurance underwriting platform must not start writing policies that will bankrupt the insurer, or refuse to write them because of illegal bias. A customer service agent must not anger customers, or give away enough refunds to make the company unprofitable. Mistakes are already common in production, because organizations seem to release AI systems without testing them enough. In McKinsey’s 2025 survey of 1,993 people, 51% of those whose organizations use AI had seen at least one negative consequence from it. Nearly one-third of all respondents had seen one that came from AI being inaccurate. This is the pattern among AI early adopters in industry this year: the software development team spends most of its time being the QA team. (Yep, it’s not just you.) It’s the reverse of what was hard about building software until very recently. The summit’s theme became show your work and the proof it asks for has to come from your own team. The 50+ sessions will stream online from October 13 to 15 and are free to attend. Three kinds of borrowed proof that don’t cover your systemThe natural shortcut is to borrow the proof: assume that the published studies, the model’s benchmark scores, or the vendor’s own validation is enough. However, each one describes something other than the system you are running. The evidence, mostly from healthcare:
Among US hospitals that use predictive models, 61% check them for accuracy on their own data and 44% check them for bias, according to |