Investment Thesis as Code

Posted by Dr. Marcel Müller | October 7, 2026


VC is not a game of finding "just good" companies, but “it is a 100% game of outliers - it’s extreme exceptions” that makes VCs successful, as Marc Andreessen [1] phrases it.

Because these outliers are so rare, a VC’s probability of catching one is directly proportional to the number of deals they process at the top of their funnel.

To filter through a sea of thousands of companies, VCs need to have a hard filter that has their written investment thesis plus the experience, learnings and in-detail specifications that a VC makes during its timeline of existence. Learnings like “we are not into AI process automation because the market is a red sea” and many other learnings a VC gets over time need to feed into this “living thesis”. Applying this as a filter to thousands of companies is challenging even with the best AI models out there. Inconsistent analyses, missed details and hidden aspects are costly and might make a VC miss the next unicorn.

My name is Marcel and I am the founder of Flowhive.VC. We build software for evaluating thousands of companies with an Investment Thesis as Code that is executable by a machine. In this article, I will break down the encoding of investment theses and how you can apply it to your pipeline today.

Table of Contents

Every VC firm has two investment theses

I am not a big fan of the overused “imagine your whole team as a bunch of AI agents and offload your work to them” narrative.

This approach oversimplifies the complexity of the process.

Let us assume for a second that you are a partner in Super VC and want to scale your team that currently consists of you and an investment analyst - let us call him Bob - to 15 analysts. The biggest problem you are going to face is not that everyone has a different bias, but that your whole collective knowledge as a VC is spread across 15 different brains in the heads of 15 different people.

Every VC has an investment thesis. Actually, VCs have at least two investment theses.

The first one is rather static and high-level, and usually lives on their website and on slides that are used to raise the funds. This first investment thesis is used to communicate about what the VC is all about to external parties of some sort. A statement you would see in an external investment thesis is “we invest in AI companies with an application to specific industries that automate a currently manual workflow” or “we invest in FinTech companies across different application areas” followed by a bunch of high-level examples. This is also usually the thesis you find on the web page of a VC.

The second investment thesis is the internal living thesis.

The living thesis feeds on the experiences that the VC eventually made. For example, in the last year, your analyst Bob has done a deep dive into 40 AI companies that automate workflows in specific industries. He found that there are a lot of drag-and-drop workflow automation platforms and looked into the market. It seems like this is a complete red sea and the market for these applications is already divided. So, there is no point in investing in yet another of these companies for Bob unless there is another bigger moat. Yet there are exceptions when you for example look into these automation builders for clinical and pharmaceutical research. But then only for areas where the company is researching highly in-demand therapeutics.

The internal living thesis consists of a lot of “yes, but” and “if not… except” and “but then…” phrases. These are complex relationships and even harder to encode.

A VC Thesis Evolves Inside a Funnel

VCs filter out the companies usually in a funnel-like structure.

On top there are a lot of companies that are being looked at relatively superficially and usually by analysts and a lot of companies are being eliminated and desk rejected right away. If Bob the analyst sees a company that is building cleaning supplies, he knows right away that this company is not in the investment scope of the VC and can discard it without having to talk to anybody else. That saves time that Bob can spend now on analyzing other companies.

The deeper a company gets in the funnel, the deeper the analysis gets and the more effort Bob and his colleagues put into the analysis of a particular company. And at the same time, the deeper a company gets in the funnel, the more people in a VC organization have a look into it.

If we want to have an encoded investment thesis that we later can automatically execute (using automation and AI), we need to make this thesis funnel-aware. That means thinking of each of the stages in the funnel and discussing what we can at that stage confidently analyze and what needs other tools or more work. For example, seeing that a company is in the broader fintech sector is pretty easy, but understanding the technology, the market size and the business pain solved needs a fundamentally deeper understanding in a later stage of the thesis.

We should not give every company a full due diligence and 2 weeks of in-depth analysis, because it would be costly and highly inefficient. Our encoded thesis needs to be applied in these stages.

How to Encode Your Investment Thesis

There are a bunch of products out there that let VCs “find companies” to “have a look at”. Products like Crunchbase, PitchBook, Dealroom, and a dozen others have built out large and helpful databases. These products all come with the option to filter by certain dimensions like the location, the general industry and the incorporation year. When we think about the fact that a VC is a step-by-step funnel, the very broad lists that you would get from products like that are completely top of the funnel and a great starting point. The data is limited so it is important not to kick out companies too early.

So your first stage of your encoded investment thesis should focus on scope. I recommend 3 broad dimensions:

  • stage of the company (pre-seed, seed, Series A,...)
  • location of the company (North America, Europe, Asia, …)
  • generalized industry (finance, manufacturing, …)

These 3 dimensions can be fairly reliably answered with data from the traditional VC data suppliers from above. For Super VC with Bob as the analyst, as an example, that is relatively easy:

  • Stage: Seed
  • Location: companies in North America and the EU.
  • Industry: finance, manufacturing, high tech, software

The first stage information in our encoded thesis includes mostly content from the static high-level investment thesis and no in-depth knowledge from your living thesis. Think of it as “everything that you can filter out with 7-8 clicks in the filter of Crunchbase”.

In the second stage of your encoded investment thesis, we are still looking into scope but also into strategic fit. And this time we do not find all the information in a single database, but we need to dig into additional information to find it. For that, I recommend you split your encoded thesis into several parts by industry and write the details of what you want and what you don’t want into a form that looks like this.

Industry: Cybersecurity for AI
Security software addressing the layers most relevant to a global brand running complex AI-agent environments - with emphasis on gaps the established Gartner-leader vendors haven't closed yet.

And then you don’t stop just at defining the industry but defining exactly what you are looking for in this industry and what you are not looking for. Both are very important especially if you want to make your thesis automatically evaluable with AI.

A) What we want in Cybersecurity companies:
  • Applications for vetting and scanning of MCP servers before they're approved for enterprise use.
  • Observability and audit logging for AI agents that ties back into existing identity systems.
  • Detection built specifically for MCP-related threats - tool poisoning, prompt injection, fake or rogue servers.
  • Enterprise-grade registries of vetted, approved AI agents and MCP servers.
  • Monitoring that flags when an AI agent acts outside its intended scope or permissions.
  • Governance frameworks that let enterprises adopt AI agents safely and at scale.
  • Visibility and control over non-human identities - API keys, tokens, service accounts, certificates - across cloud, SaaS, and on-prem environments.

B) What we don’t want in Cybersecurity companies:
  • no categories already saturated by Gartner leaders (CSPM, source-code security, general vulnerability scanning)
  • no offensive security, defense, or cyber-warfare tools
  • no consumer-only security products
  • no OT or production-network security
  • no homomorphic encryption or similar infrastructure

Let us examine these examples for a bit.

We define what we are looking for in a company and exactly what we are not looking for. Paying close attention, you see that the non-desired criteria are much broader (see “no offensive security, defense, or cyber-warfare tools”) than what we desire (see “Monitoring that flags when an AI agent acts outside its intended scope or permissions.”). That is an optimization done already for AI agents to be able to execute well. LLMs are naturally good with language, but it is hard for them to have the understanding of the in-depth mechanics and synergies, together with the current state of the market, if you do not define them.

All these very in-depth criteria come from your living thesis. So whenever you accept or reject a company, try to understand what you can learn from it and what is generalizable in a way that we can fit it in the living thesis document.

And this principle is applied in the same way in the other stages in the VC funnel. You can make a stage of whatever you like, for example quality criteria, traction criteria and even go so far down as to do risk and threat modeling. As long as you are able to encode it, it can be used by a machine.

Executing an Encoded Thesis with AI Remains a Challenge

The goal of the whole process of encoding your thesis is to make it evaluable with AI in a way that you can process thousands of companies with the same quality of results. That’s where you need a more complex AI orchestration than “hey Claude Code, here is my investment thesis, can you find me startups?” to solve the following challenges:

Normal LLMs with a generalized but not specialized orchestration VC harness around them have a limit in-depth of web search (when you are trying to find sources and evidence of a company). When we are looking a bit closer into the inner workings of “standard LLM-based agents” that you get today with Claude, ChatGPT, Gemini, Grok and others out of the box, we see that the number of web pages that an agent will look at when using their web search tools is limited in most standard cases to about 20-60 web pages. Deep research functionality that is provided, for example, by Google’s Gemini can go deeper but none of these models comes out of the box with a configuration that screens hundreds of web pages deep. There are today specialized web research models that are good at digging up information but not good at analyzing information. So getting the most out of a really deep background search for a company requires not a single but many search prompts focusing on different aspects and a separate model for analysis and distillation.

This also means, if you really want to do this at scale and not do “a quick ad-hoc prompt and call it a day”, you need different models and tools stitched together to get the maximum of results. Today, every other day there is a new model published that is better, usually at a very specific focus point. Staying on top of these models is also a challenge in itself.

Once you have the best models and your background research and your search model has found all the materials to judge a company along its criteria, for example, LinkedIn accounts and account posts to judge how well the founders are suited, or 30+ pages of whitepapers explaining the moat, the best practice is to look into every dimension of your judgment individually. If you gave your whole corpus of research data to a model and say “that’s the startup, now tell me if I should invest”, you will not get complete or reproducible results and in most cases, even worse, you will get false results. This is for the reason of context length. Modern LLMs can generally process up to 1M tokens. Yet, just because they can process it, it does not mean they are good at identifying the exact needed information to judge exactly one dimension that is relevant under a huge pile of unstructured data. This is called the needle in the haystack problem and often is the cause of very bad results.

And once you have mastered all of this, you come to your final opponent: costs. If you want to throw your most capable, most expensive model on all of your tens of thousands of companies, you are not in the area of paying pennies for the evaluation of a single company anymore. We are talking about significant costs here for each company and that multiplied by your thousands of companies you want to evaluate… This adds up quickly.

In my day-to-day life I work at my VCtech startup FlowHive.vc where we build such an orchestration engine to execute evaluations of startups at scale. We have tweaked these evaluation calls, selected the best models (and continuously updated them) and tweaked them to produce reproducible costs and reproducible quality. This is a task that first seems trivial, because out of the box every one of these models gives you something that “does not look that bad”. But getting it right in a serious environment is not easy.  

How good is an Encoded Thesis?

With the tools from above you are able to over time create your own encoded investment thesis and evaluate hundreds or thousands of companies. But how can you be sure that your automation works at a high-level of quality? 

The method I have used in the past is to treat it as a statistical evaluation.

Unless you are a completely new fund, you have looked into hundreds of companies and ideally recorded which ones you kicked out, which ones you advanced deeper into analysis and which ones you invested in. Your notes on the decisions helped you to build the encoded thesis. We take your manually analyzed companies and the decisions as the basis for evaluation and we try to see how well our encoded thesis is able to reproduce the same (or better) results as humans would have done.

Let us assume we have evaluated a company called Cyber Company AI. Bob, our human analyst, evaluated it in his stage 1 analysis as up. If our encoded thesis AI analysis also marks it as up, we are all good. In the happy case, we call this a true positive. However, if our AI automation had said the company is out while Bob said up, that would be a false negative and this is costly. So imagine: you have invested in a company but the AI that you are trying to teach with your encoded thesis how to evaluate startups like you would evaluate them would reject the company. This is the absolute worst case because it could be a missed unicorn or a fund returner that we urgently need.

If our AI automation says the company is interesting but we as humans in the beginning discarded it already (false positives), we do not cause a lot of damage, just not saving the human a lot of time. And finally, true negatives are again the happy case.

When we adjust our encoded thesis, we want to make sure to minimize false negatives (not kicking out companies that the human evaluators would have advanced) while trying to maximize true positives. This metric is called sensitivity (or recall) and measures how good we are at not missing what humans would have advanced. The ratio of true negatives to all companies that were labeled as negative (by a human) is our specificity and signifies how much time we saved. Usually the forces of false negatives and false positives work in different directions. Of course you could have 0% false negatives when you just have your thesis encoding so broad that it lets everything through. But then you have the same capacity issue as in the beginning.

From past cases that I worked on at FlowHive.vc where we evaluated for our VC customers, we usually reach a recall in the high 90s (usually around 97%) and specificity around 70% (signifying the time saved) and an overall accuracy of 90%. That is possible but requires a very in-depth tuning of the investment thesis.

Conclusion: Scaling, Reliability and Quality

With an encoded thesis and an automated pipeline, you have a lot of advantages. You can now scale beyond your current team size, your thesis and decisions are reproducible and you have aligned quality regardless of who evaluates a company.

This method is helpful for large VCs who already have a big pipeline and want to gain some speed, but also for smaller investors and angels who would like to operate at a higher scale than they have capacity for.

Less time wasted, more time spent where it should be spent. And in the end, the whole VC ecosystem benefits from it.

About the Author

Dr. Marcel Müller

Marcel is an AI specialist and entrepreneur. He has a background in enterprise systems, process automation and analytics. He leads Jaden Data, the company that builds www.flowhive.vc as an orchestration platform behind encoded investment theses.

[1] Andreessen, M., & Conway, R. (2014). "Lecture 9: How to Raise Money." CS183B: How to Start a Startup. Stanford University. Transcript and video available via Y Combinator.

You might also enjoy

Startup Valuation: What Founders Are Actually Seeing in 2026
Startup Valuation: What Founders Are Actually Seeing in 2026

This article works through the numbers layer by layer:what founders are raising at by stage and sector in 2026.

By Colin McCrea| September 26, 2026
The 23 Types of Startup Investors & How to Approach Each
The 23 Types of Startup Investors & How to Approach Each

Learn the different investor types beyond your typical VCs, and how to fundraise from them.

By Stéphane Nasser| September 9, 2026
How FOMO.ai Raised $2.2M To Replace Marketing Agencies with AI
How FOMO.ai Raised $2.2M To Replace Marketing Agencies with AI

An interview with Dax Hamman, co-founder of FOMO.ai, on raising entirely through SAFE notes, why getting fast rejections is a strategy, and the concept of building gravitational pull around your company.

By Shaun Gold| August 11, 2026