AI is rapidly reducing the cost of turning an idea into working software. That doesn't mean difficult software engineering has disappeared. Production systems still require judgment, reliability, taste, security and care. But as implementation becomes cheaper at the margin, knowing what to build becomes relatively more valuable.
That shifts more strategic weight onto problem discovery, market discovery, research and strategic search. It also raises a broader question: can startups and smaller companies use AI to acquire organizational capabilities that historically required dedicated departments inside much larger organizations?
Strategy and research are interesting examples because they are not mainly about producing more output. They are about deciding which output is worth producing at all. Models can search and synthesize information quickly, reason about hundreds of possibilities at low marginal cost, and help preserve the results of that work.
So I started an experiment: could I build something resembling a continuously operating AI strategy and research function for a startup?
The first versions sounded quite sophisticated. Every day the system collected news and market signals. AI analysed them against what we were building at TailyX. Different processes generated strategic ideas. Claude and ChatGPT could challenge ideas independently. Ideas were scored, compared and ranked. The strongest appeared in an email every evening. I eventually had processes for strategic thinking, organic idea generation, competing models and idea-quality evaluation.
There was just one problem. I wasn't particularly impressed by the ideas.
They weren't stupid. That was what made the problem interesting. They were usually perfectly sensible. But again and again they clustered around things we were already doing: lead qualification, websites, attribution, intake, AI agents, APIs, WhatsApp, sales workflows.
The system was becoming better at telling me how to extend the world I already occupied. It wasn't getting much better at discovering worlds I hadn't considered.
When I examined how the system actually worked, the reason became obvious. We were rewarding things like:
Those are all rational things to measure. They're also an excellent recipe for convergence.
They also concealed a deeper problem: there is no objectively best strategic idea without specifying what the system is optimizing for.
The best search for the fastest credible route to substantial revenue is not the same as the best search for discovering a genuinely novel venture-scale company. Both are different again from the best search for the strongest fit with TailyX's existing assets, customers, distribution and founder capabilities. Selecting the objective function is itself a strategic decision.
My system had made that decision implicitly. It rewarded fit, evidence and adjacency, so it naturally converged around TailyX. Suppose an existing thesis accumulates evidence. Tomorrow's AI research finds another development supporting it. Its score increases. The system searches around it again. More supporting evidence appears. And gradually: existing thesis → nearby search → supporting evidence → higher score → even more nearby search.
The AI wasn't failing. It was optimizing exactly what I'd asked it to optimize. The problem was that genuinely novel ideas start with an enormous disadvantage. A new idea has no accumulated evidence. It has no previous ranking. It may have no relevant customer conversations. It may not resemble our existing product. It may require technology we haven't built. The stranger the idea, the worse it initially performs against an evaluation system designed around what we already know.
That led to the most important realization from the experiment so far:
The breakthrough wasn't a better prompt or a better model. It was realizing I had built an optimizer when what I needed was a search engine.
My first instinct was straightforward. If 20 ideas aren't producing anything extraordinary, generate 100. Or 1,000. Or perhaps 10,000.
But there's another problem. If all 10,000 are generated from roughly the same conceptual starting point, you've just searched the same neighborhood more thoroughly.
So I've started changing the unit being searched. Instead of asking AI "what company should I build?", the system first searches for much smaller hypotheses:
CHANGE → ACTOR → PROBLEM/OPPORTUNITY → ECONOMIC CONSEQUENCE
For example: a technology cost collapses. Who does that affect? What previously uneconomic activity becomes possible? Whose business model changes? Where does authority move? What new transaction becomes possible? What expensive human workaround stops making sense?
Only much later should the system ask what company could exist because of that change. This creates a very different process:
world changes → thousands of hypotheses → deduplication → diversity selection → evidence → opportunity clusters → recombination → company theses → adversarial evaluation
The email is almost the least important part.
I also realized I had been asking one AI system to answer two fundamentally different questions.
The first is: how can TailyX win? That's an exploitation problem. Existing technology matters. Current customers matter. Distribution matters. Evidence matters. Proximity matters. I absolutely want AI searching aggressively around those things.
But there's another question: what important company should exist because the world is changing? That's an exploration problem. If I tell that system everything about TailyX before it starts searching, TailyX becomes a gravitational field. Suddenly everything looks like another lead product, website product, sales agent or marketing tool.
So I'm now separating the two. One engine deliberately searches opportunities adjacent to TailyX. Another starts much further away. During its early search it should effectively be blind to what TailyX currently does. It searches technological discontinuities, changing economics, demographics, regulation, robotics, energy, ownership, infrastructure, computing and other shifts.
Only after promising opportunities survive should the system ask: does this founder have an unusual right to win here? That ordering matters.
This evolution has also changed how I think about "agentic AI." An agentic strategy system isn't particularly interesting if it means giving one AI agent a giant prompt saying "be my Chief Strategy Officer."
The architecture I'm finding more interesting has specialization. One process asks: what changed today? Its job is evidence collection, not company generation. Another asks: what opportunities adjacent to the existing company follow from these changes? Another deliberately searches outside the company's existing territory. Other processes can eventually perform semantic deduplication, novelty detection, evidence verification, hostile evaluation and recombination.
And importantly, agents should be allowed to return: nothing passed the bar today.
That's quite different from asking a chatbot for five ideas. It starts looking more like a collection of specialized research workers operating against a common knowledge base. The eventual system I imagine isn't an AI that "knows strategy." It's a machine that continuously searches a strategic possibility space.
One experiment deliberately exposed the AI to unrelated domains and forced it to extract mechanisms that might transfer elsewhere. The point wasn't to tell the AI to "be more creative." It was to manufacture conceptual distance.
That experiment also taught me something about evaluation. If I tell the generator exactly what the novelty evaluator wants, I risk getting ideas engineered to look novel. So generation and evaluation sometimes need to remain separate. I'm even withholding intermediate novelty judgments from myself during one experiment because seeing them could influence my later assessment.
That's starting to make this feel less like prompt engineering and more like experimental design.
Not yet. And that's an important qualification.
The current system cannot replace great customer conversations, proprietary industry knowledge, experienced strategists or the intuition that comes from operating inside a market. Nor have I yet demonstrated that searching thousands of hypotheses produces better companies. That's the experiment.
But I think it is already better than the alternative available to many small companies: not doing systematic strategic research at all. A startup doesn't necessarily need a dedicated analyst team to monitor hundreds of markets if it can deploy specialized AI processes that search continuously, preserve what they learn, challenge one another and escalate only the small number of things that deserve human attention.
The founder's job then changes. Instead of generating every possibility personally, the founder increasingly defines what should be searched, what constitutes evidence, what should be rejected, where exploration should occur, and what deserves scarce human attention.
The missing piece is memory. At the moment, AI systems are still too prone to behaving as though each research cycle starts from scratch.
What I ultimately want is a cumulative hypothesis store. Every hypothesis should retain its evidence, semantic relationships, previous evaluations, rejection reasons and subsequent developments. Then the system can know: we considered this six weeks ago; we rejected it because X; today Y changed; therefore it deserves reconsideration.
It should also learn from my decisions. If I repeatedly reject ideas because they're merely AI wrappers, that is information. If I repeatedly investigate businesses involving a new economic control point, that's information too. Eventually there may be enough labelled decisions to introduce learning-to-rank or other ML techniques alongside the language models.
At that point the system becomes less like "ask ChatGPT for strategy," and more like: search → learn → remember → allocate research → challenge → experiment → update. That's much closer to what I think an agentic strategy capability for a small company could eventually become.
The question is no longer only whether a startup can produce more output. It is whether it can build enough strategic search capacity to decide which output is worth producing at all.
I don't yet know whether this experiment will consistently discover exceptional opportunities. But it has already changed the question I'm asking.
I started with: "how do I get ChatGPT and Claude to give me better ideas?" I'm now asking: "how do I design a machine that searches enough of the possibility space that something genuinely non-obvious has a chance to emerge?"
Those are very different problems. And I suspect the second question is where things start getting interesting.
Related: for the production system this research runs alongside, see AI lead qualification software.