moki

When to kill a startup idea: the five gates that come before any score

Kill a startup idea when one of two things is true. A must-pass condition clearly fails, in which case no score matters and nothing downstream can rescue it. Or your honest range of outcomes tops out below your own bar, meaning even the optimistic end is not worth the next year of your life.

Everything else is a hold. That word does real work here, because a hold isn't a polite kill: it means the decision is blocked on something you can go and find out.

That's the shape of the decision. The rest of this post is the machinery: the five must-pass gates our audit engine runs before it computes anything, the two rules that catch what an average hides, and how to run all of it by hand on a Saturday. These are the actual criteria our audit engine runs, thresholds included, not a generic listicle.

Generating ideas got cheap. Killing them did not.

The supply of ideas isn't your problem any more. A model will hand you ten before lunch, each with a plausible market description attached, and the tools built to do exactly that are multiplying. We pulled Google Trends the morning this went up: among the fastest-rising related queries for "startup idea", two were product names, ideabrowser up 24,100% and ideaproof up 11,100% year over year.

What none of them hand you is permission to stop.

That asymmetry is the whole reason kill criteria matter now. When ideas were expensive to produce, the scarce skill was having one. Now that building collapsed to a weekend and generating collapsed to a prompt, the scarce skill is throwing them away on purpose.

You cannot make this call afterwards

Here's the part that catches everyone. The moment you start building, you become the worst possible judge of whether to continue.

That isn't a character flaw, it's a documented and stubborn bias. In the classic 1985 study on sunk cost, Hal Arkes and Catherine Blumer ran a field experiment on theatre season tickets: customers who happened to pay full price attended more plays over the following six months than customers randomly given a discount. Same shows, same value ahead of them, different amount already spent. The money was gone either way, and it changed their behaviour anyway.

Now scale that from a theatre ticket to eight months of nights and weekends, a name you love, and a demo your friends called cool.

The fix is procedural, not emotional. You write the conditions that would kill the idea before you have anything invested in the answer, with a number and a date, while you are still capable of being fair. After that the decision is a checkbox rather than a heartbreak. This is also the reason a chatbot cannot make the call for you: it validates you, not your idea, and by the time you ask, you have already told it how much you want this.

The five gates

Our engine runs five knockout gates before it computes a single score. They are must-pass. If any one clearly fails, the audit stops and returns a kill, because no amount of strength elsewhere fixes a structural impossibility. A brilliant product nobody may legally sell isn't an 82 out of 100.

The judging rule matters as much as the list, so here it is verbatim from the spec: when unsure, prefer null over false, but if a fatal risk is plausible yet unverified, prefer null over true. Do not kill on speculation. Do not wave a risk through either. An unresolved gate is not a pass.

Run them in this order, because they get more expensive to check as you go down.

  1. Legal and ethical feasibility. Can this legally exist, sold by you, to these people, in this market?

    Self-test, about 30 minutes (our judgment, not data): search your category plus your country plus "regulation", "banned", or "licence". You're not looking for a legal opinion, you're looking for whether a licence, a registration, or a hard prohibition stands between you and your first customer. A licence you can get is a cost. A prohibition is a wall.

  2. Minimum viable market. Is there enough demand to justify the thing existing at all?

    Self-test, about 30 minutes: read the demand curve and the modifier layer for your category. Our method for reading market demand off Google is the long version. The failure you are looking for is not "crowded", it is "empty and flat": nobody searching, nobody complaining, nobody selling. Crowded with rising demand is a validated market. Empty is a graveyard, and silence is the most misread signal in validation.

  3. Technical feasibility. Can this actually be built, by a team like yours, with today's technology?

    Self-test, about 20 minutes: name the one component you have never built before and cannot buy. If the answer requires a research breakthrough, a dataset you cannot obtain, or a permission a platform does not grant, that is the gate. Note that "an API might remove this capability" belongs in gate 4, not here. Whether a model update makes you redundant is a different test entirely.

  4. Strategic alignment. Does this fit what you are actually trying to do, and does it survive its dependencies?

    Self-test, about 15 minutes: write one sentence saying why you specifically should build this, and one naming the platform or partner who could end it with a policy change. If the first sentence is empty, you will quit in month seven for reasons that have nothing to do with the market.

  5. Resource availability. Do you have the money, time, and access to reach a first paying customer?

    Self-test, about 15 minutes: count the months of runway you would actually commit and the hours per week you will actually work, not the ones you would like to. Then ask whether the first customer is reachable inside that budget. This gate kills more good ideas than bad markets do, and founders skip it because it is about them rather than the idea. If that budget only works once you leave your job, set the exit threshold before you decide.

Three answers, not two

Most kill frameworks give you pass or fail, which forces you to lie. A seed-stage founder honestly doesn't know their regulatory exposure yet, and being made to choose between "it's fine" and "kill my idea" produces a false attestation almost every time.

So our gate cards take a third answer: I don't know yet.

It doesn't kill the run, and it isn't a free pass either. It prices the ignorance. An unresolved gate caps confidence at 0.6 on the dimension it maps to, no matter how strong the outside evidence looks, and that cap binds your own override too. Lower confidence widens the band. A wider band makes a Make verdict harder to reach without making it impossible, which was a deliberate call: we price uncertainty, we don't punish honesty.

The report then turns every unresolved gate into a "learn this first" item in your roadmap. That's the honest output when the truth is that you don't know yet.

Two rules that catch what an average hides

Pass all five gates and you get scored across six weighted dimensions: market attractiveness (0.22), differentiation (0.20), financial reward (0.18), strategic leverage (0.15), risk (0.13), and team and resource fit (0.12). Each dimension returns a point and a confidence, and confidence becomes the width of the band rather than a decoration.

Two rules then decide the verdict, and both exist because a weighted average is a liar.

The veto rule. Any single dimension scoring 2 or below out of 10, held with confidence 0.7 or higher, blocks a Make outright regardless of the total. A confident disaster in one place is not something a strong average gets to average away.

The band rule. The verdict reads the ends of the range, never the midpoint.

VerdictConditionWhat it means
MakeBand's low end ≥ 65, and no vetoEven the pessimistic case clears the bar
KillBand's high end ≤ 40Even the optimistic case does not
HoldAnything else; band width > 12 means "reduce uncertainty"The idea is not the blocker, your information is

Read the kill row again, because it's the part most people get backwards. We only return a kill when the best case is still bad. That's a deliberately hard bar, and it means most bad news arrives as a hold with a leaning rather than as an execution.

A range with a stated confidence, never a fake 73 out of 100.

What it looks like when nothing fails

Here is a real audit, public and unlocked: LedgerAgent, bookkeeping software for independent US real-estate agents.

All five gates passed. Legal, market, technical, strategic, resources: clean. Under the old pass-or-fail logic that reads like a green light.

The verdict was Hold. Differentiation scored 4.0 out of 10, because the evidence layer found purpose-built competitors already advertising the exact features the founder described as the wedge. The report calls those claimed differentiators "not a novel wedge but table-stakes functionality already advertised by incumbents".

That's the case the five gates alone can't catch, and it is by far the most common one. Ideas rarely die of illegality or impossibility. They die of being fine.

The cost of a false kill

A framework that only ever says kill is as useless as a chatbot that only ever says build. Both replace judgment with a reflex.

The failure statistics are worth reading in that spirit. BLS data shows 34.7% of establishments born in 2013 were still operating ten years later, which is a long way from the "90% fail" folklore. CB Insights, analysing 431 VC-backed shutdowns since 2023 and publishing in March 2026, found 70% ran out of capital and 43% had poor product-market fit underneath. Companies mostly do not die of one dramatic flaw. They die of a slow mismatch nobody wrote down early enough to notice.

So the goal isn't a high kill count. The goal is that the decision gets made by a condition you set in advance rather than by exhaustion, a bad week, or the last person whose opinion you asked.

A kill you cannot justify in writing is just a mood.

Run your idea through the gates

Write your five gates tonight, with a number and a date on each. If you would rather see them applied to your specific idea, with live Google demand data and an adversarial kill case attached, that is exactly what our audit does. It runs the five gates, scores six dimensions in honest ranges, and is allowed to tell you to stop.

Score an idea — free

FAQ

Is it true that 90% of startups fail?

There is no authoritative source for the 90% figure, and the real numbers depend entirely on what you count. US Bureau of Labor Statistics data shows 34.7% of business establishments born in 2013 were still operating in 2023, so roughly two thirds closed over a decade rather than nine in ten quickly. Venture-backed startups are a narrower and riskier population: CB Insights analysed 431 VC-backed companies that shut down since 2023 and published the results in March 2026, finding 70% cited running out of capital as the final cause and 43% poor product-market fit underneath it. Treat 90% as folklore, and treat your own idea as a specific case with specific failure conditions you can write down in advance.

What is the #1 mistake startups can make?

Building something the market does not need. In CB Insights' March 2026 analysis of 431 VC-backed shutdowns, poor product-market fit sat behind 43% of failures, while running out of capital (70%) was usually the last event rather than the root cause. There is a quieter version of the same mistake that costs less money but more time: never writing down what would prove the idea wrong. If you have no falsifiable condition, you cannot fail a test, so you keep going on feel until the money decides for you.

What is the 6 month rule in business?

It usually refers to keeping six months of operating expenses in reserve, and it is a liquidity guideline rather than a validation rule. It says nothing about whether the idea is any good. If you want a runway rule that actually informs a kill decision, invert it: decide today which specific evidence you must see before you spend the next six months of runway, and write down the date you will check. A reserve protects you from a cash surprise. A dated evidence threshold protects you from spending the reserve on the wrong idea.

What is the 80/20 rule for startups?

It is the Pareto principle applied loosely to startups: roughly 80% of your results come from about 20% of your effort, customers, or features. It has no single authoritative definition and no agreed measurement, which makes it useful as a focus prompt and useless as a kill criterion. You cannot fail an 80/20 test, because nothing about it is falsifiable. Kill criteria need the opposite property: a specific condition, a number you set in advance, and a date by which you check it.

Get the honest read on your own idea.

A reasoned Make / Hold / Kill in minutes — free, no signup.

Score an idea — free →

← All notes