The news is full of sobering data on enterprise AI. A recent MIT study found that 95% of companies get zero return on their AI investments. For economists, this looks like a bubble ready to burst. But for those of us on the ground, in the trenches of software development, these studies reflect a broad average perspective where, while accurate, the reality can be far more nuanced.
So here is a techie SEMs take on why you should still be experimenting with AI, despite the undertones of foreshadowing, in and among all the hype. I've spent the larger part of my life in these "decade's old code bases" with high regulation and ridiculous complexity. By all accounts, this environment should present the exact kind of "brownfield" challenge where, as a recent Stanford study of over 100,000 developers found, the productivity gains from AI face diminishing returns. This is true, it becomes exponentially harder as you lean in with AI on these brownfield environments, this is a tough place to look for quick gains. It requires advanced prompting, agent orchestration, and MCPs to get even close to production-grade code. Yet, my experience with Large Language Models (LLMs) tells a different tale.

The progress has been staggering. Just a few months ago, AI capabilities were significantly limited compared to what LLMs and complementing toolsets like GitHub Copilot and Claude Code are able to do this year. Today,my teams are regularly generating code that satisfies 90% of a requirement, not just for simple tasks but for production-ready, business-critical logic. And it's not "garbage in, garbage out" code; it's relatively clean, well-structured output that conforms to our existing standards and passed borderline overbearing quality checks. This isn't theoretical. It’s a reality we're working with every day.
I understand the skepticism. It took me a long time to get AI-generated code working even on my own pet projects, let alone on enterprise-grade software. While this process can be challenging, expecting quick returns might not be a reasonable expectation. Of course, most AI projects stall in the pilot phase, as shown by the MIT study, but that's to be expected. Given the infancy of this new generation of models, it stands to reason that most companies playing with it will inevitably fail—that's the nature of emerging technology. We're still in the pilot stage, and not enough time has passed for anyone other than Jensen Huang and Sam Altman to be getting serious "ROAI", at least not at the application layer. ChatGPT was only released in late 2022, with an enterprise version coming in late 2023. Most companies have yet to properly roll that out, let alone start to play with specialized agents.
Despite all this, AI is still a net positive, and all signs indicate these models are only getting better. Early adoption is not for everyone, and context and budget matter. Luckily, you don't have to go all in. Consider GitHub Copilot or Claude Code, especially for small teams of six or more engineers. At roughly $2,500 per annum, most SMEs should be able to afford that. At that price, most enterprises just don't have a good excuse not to be trying to operationalize AI agents. The teams that will win are those bold enough to try the things the LLM can't do now, because in six months' time that capability will probably become available. By then, you'll be miles ahead in having built that muscle compared to the late adopters and your competitors. While ROI with AI might be slow to realize, in the end, winning may come down to those who can work with AI and those who can't.
