# Startup hypothesis examples: what to test, and what a good hypothesis looks like

Source: https://hypothis.ai/blog/startup-hypothesis-examples
Published: 2026-10-08
Updated: 2026-10-09
Author: Kapil Hingu

Every startup is a stack of guesses. Someone has this problem, it hurts enough to act on, they'd switch from what they use now, they'd pay this much, and you can reach them. A startup hypothesis is one of those guesses, written down precisely enough that real evidence could prove it wrong.

In science, a hypothesis is a proposed explanation that can be tested and falsified. Startups borrow the idea for the same reason: you can't learn anything from a claim that can't fail. "People want this" survives every result. "Freelance designers lose two hours a month chasing late invoices" can come back false, which is why a round of research can teach you something about it.

Below are the areas every founder should form hypotheses about, weak and strong examples for each, and how to validate or invalidate them with real customers. For the mechanics of writing a single hypothesis, start with [how to write a startup hypothesis](/blog/writing-testable-hypotheses).

## What a good startup hypothesis looks like

A hypothesis is ready to test when it passes all six of these checks.

| Check | Weak | Good |
|---|---|---|
| Specific segment | "Small businesses" | "Independent bakeries with 2-10 staff" |
| Observable behavior | "They struggle with scheduling" | "The owner rebuilds the staff rota by hand every week" |
| Magnitude | "It's a pain" | "It takes 3+ hours a week" |
| Falsifiable | "Some would find it useful" | "At least 6 of 10 qualified owners describe this unprompted" |
| Tied to a decision | "Interesting to know" | "If false, we drop bakeries as the first segment" |
| Testable by you, now | "Bakeries in 40 countries" | "Bakeries you can reach in two weeks" |

Founders skip the fifth check most often. If a hypothesis wouldn't change what you do whichever way it comes out, it's trivia, and you can cut it.

### A four-line template

Write every hypothesis in four lines:

> **We believe** [specific segment] [does / experiences / would pay for something specific].
> **We'll know we're right if** [observable signal] reaches [threshold].
> **We'll know we're wrong if** [the kill condition].
> **If we're wrong, we will** [the decision this changes].

The second and third lines are your validation criteria, and you write them *before* you collect a single response. Once data starts coming in, every founder is tempted to reread a weak result as "promising." Thresholds you committed to in advance keep you from moving the goalposts.

## The areas every startup should form hypotheses about

Most ideas rest on the same handful of load-bearing claims. Test them separately and you learn *which part* of the idea is working, which decides what you do next. "Real problem, wrong price" calls for a pricing experiment. "No real problem" is a reason to stop.

### 1. Problem hypothesis

**The claim:** a specific group has a specific problem, and it hurts enough that they already spend time, money, or effort on it.

- **Weak:** "Remote teams struggle with communication."
- **Strong:** "Engineering managers at 20-100 person remote startups spend 2+ hours a week chasing async status updates, and have tried at least one tool or ritual to fix it in the last year."

**How to test it:** ask about the last time it happened instead of asking whether it's a problem. "Walk me through the last time you needed a status update from your team" beats "Is async communication hard?" every time. Listen for whether they bring up the problem on their own.

**Validated if:** most qualified respondents describe the problem unprompted, with a recent, concrete example and evidence of spend or a workaround.
**Invalidated if:** the problem only appears when you suggest it, or people agree it exists but have never done anything about it.

The problem hypothesis is almost always the riskiest, so test it first. If the problem isn't real, nothing else on this list matters.

### 2. Customer segment hypothesis

**The claim:** *this* group, defined narrowly, feels the problem most acutely and is the right place to start.

- **Weak:** "Our customers are freelancers."
- **Strong:** "Freelance developers billing over $80/hour feel this more than designers or writers, because their invoices are larger and late payments hurt cash flow more."

**How to test it:** recruit two or three adjacent segments into the same round and compare them. The segment you assumed was best often comes second. This is also why screeners matter: a response from the wrong person is noise, and it dilutes the answers that count.

**Validated if:** one segment shows clearly stronger pain, frequency, or spend than the others.
**Invalidated if:** the pain is spread thinly across everyone, which usually means it's mild for everyone.

### 3. Behavior hypothesis

**The claim:** people currently handle the problem in a specific way, and would change that behavior for something better.

- **Weak:** "People would use an app for this."
- **Strong:** "Most target users currently track this in a spreadsheet they update weekly, and at least a third have tried and abandoned a dedicated tool."

**How to test it:** ask what they do today, step by step, and what they've tried before. Past behavior is the most honest signal you can collect, and stated future behavior ("I'd definitely use that") is the least.

**Validated if:** there's a clear existing workaround that costs something real, and people have already shown they'll try alternatives.
**Invalidated if:** the current workaround is good enough and nobody has looked for something better. Your real competitor is often a spreadsheet and a shrug.

### 4. Willingness-to-pay and pricing hypothesis

**The claim:** the segment would pay a specific price, and that price makes sense against what they already spend on the problem.

- **Weak:** "People would pay for this."
- **Strong:** "Independent accountants with 30+ clients would pay $40/month, because they already spend roughly that on two partial tools that each solve half the problem."

**How to test it:** anchor on current spend before you mention any price. Ask "What do you pay today to deal with this, in money or hours?" and then test price ranges instead of a yes/no. Compliments cost nothing; a pre-order or a signed letter of intent is the strongest signal you can get before building. Question wording is covered in [willingness to pay survey questions](/blog/willingness-to-pay-survey-questions).

**Validated if:** current spend on workarounds is at or above your price, and a meaningful share accept the range without heavy hedging.
**Invalidated if:** people love the idea but currently spend nothing on the problem. That's the most common way a "great idea" turns out to be a hobby.

### 5. Solution-fit hypothesis

**The claim:** *your* specific approach solves the problem better than what people use now, by enough to justify switching.

- **Weak:** "Our AI-powered platform will help users be more productive."
- **Strong:** "Showing late invoices in the tool freelancers already use for time tracking will beat a standalone invoicing app, because they won't adopt a second dashboard."

**How to test it:** describe the approach plainly, without a pitch, and ask how it compares to what they do now and what would stop them switching. Pay attention to the objections; they tell you more than the enthusiasm does.

**Validated if:** people can say specifically what they'd stop doing if they had your solution.
**Invalidated if:** reactions are positive but vague ("cool, I'd try it") and nobody names a concrete thing it would replace.

Test solution fit after the problem is confirmed. When you ask people to judge a solution to a problem they don't have, you get polite noise.

### 6. Market hypothesis

**The claim:** the segment is large enough to matter, and you can find and reach enough of them.

- **Weak:** "The market for project management software is $7B."
- **Strong:** "There are at least 15,000 independent physiotherapy clinics in the US and UK, and they cluster in two professional associations and one active subreddit we can reach directly."

**How to test it:** count from the bottom up. Tally the actual reachable units (companies, people, communities) instead of quoting an industry report. Then look at how hard it was to recruit your research round, which gives you an early read on your future acquisition cost.

**Validated if:** you can count the segment from real sources, and recruiting qualified respondents was feasible.
**Invalidated if:** you struggled to find even 15 qualified people. If research recruiting is that hard, customer acquisition will be harder.

### 7. Channel hypothesis

**The claim:** you can reach this segment through a specific channel at a cost the price supports.

- **Weak:** "We'll grow through social media."
- **Strong:** "Shopify store owners doing $10k-$100k/month can be reached through two app-store categories and three Facebook groups, and at least 5% of a cold group post's viewers will click through."

**How to test it:** treat your recruiting effort as a channel test. Track where qualified respondents actually came from, and what it took to get them.

**Validated if:** at least one channel produced qualified people at a reasonable effort.
**Invalidated if:** every qualified respondent came from your personal network. Friends doing you a favor won't scale into a customer base.

### Hypotheses you can't test yet

Retention ("they'll still use it after three months"), referral ("they'll tell colleagues"), and unit economics need a working product and real usage before you can test them properly. Write them down so you don't forget them, and don't expect a pre-build survey to validate them.

## Which hypothesis to test first

You won't have time to test everything exhaustively in round one, so rank hypotheses on two axes:

- **Impact:** if this is false, how much of the idea collapses?
- **Uncertainty:** how little real evidence do you have right now?

Start with the hypothesis that scores high on both. For most ideas, the order looks like this:

1. **Problem.** If it's not real, stop.
2. **Customer segment**, so every later round talks to the right people.
3. **Behavior and willingness to pay** together, because current spend connects them.
4. **Solution fit**, once you know the problem is worth solving.
5. **Market and channel**, which rounds one to three partly answer through how recruiting went.

In this order, a bad first round kills the idea cheaply, in weeks, instead of after five rounds and a prototype. [When to kill your startup idea](/blog/when-to-kill-your-startup-idea) covers what that decision looks like.

## How to validate or invalidate a hypothesis

### Match the method to the claim

| Hypothesis | Best evidence | Weaker evidence |
|---|---|---|
| Problem | Unprompted descriptions of recent incidents | "Yes, that's a problem" when asked |
| Customer segment | Side-by-side comparison of segments | One segment, no comparison |
| Behavior | Existing workarounds and past tool switches | "I would use that" |
| Willingness to pay | Current spend, pre-orders, letters of intent | "I'd pay for that" |
| Solution fit | Specific things they'd stop doing | General enthusiasm |
| Market | Bottom-up counts and recruiting difficulty | Top-down industry reports |
| Channel | Where qualified respondents actually came from | Where you plan to advertise |

In every row, past behavior beats future intent and specifics beat sentiment. That's the core of [the Mom Test](/blog/mom-test-for-solo-founders), and it applies to surveys as much as to interviews.

### Score each hypothesis

After a round, give every hypothesis one of three outcomes:

- **Confirmed:** the evidence hit the threshold you wrote down.
- **Weakened:** the evidence pointed against it, or hit the kill condition.
- **Inconclusive:** not enough qualified responses, or mixed signals. That's a legitimate result, and it means you test again with a sharper question.

You don't need statistics at this stage. With 10 to 20 qualified respondents you're looking for a clear pattern, not significance, and "nine of twelve described the problem unprompted" is a strong signal. For sizing, see [how many customer interviews you need](/blog/how-many-customer-interviews).

### What a weakened hypothesis tells you

A weakened hypothesis is cheap information about which part of the idea to change. A weakened pricing hypothesis with a confirmed problem means you keep going and change the model. A weakened problem hypothesis means you stop, and you've saved the weeks you would have spent building.

## A worked example: one idea, a full hypothesis set

**The idea:** inventory sync for small sellers who sell the same products on both Etsy and Shopify.

- **Problem:** Sellers with 50+ SKUs on both Etsy and Shopify oversell at least once a month because stock counts drift between platforms. *Wrong if:* fewer than half of qualified sellers report an oversell in the last 60 days.
- **Segment:** Handmade sellers with 50-500 SKUs feel this more than print-on-demand sellers, whose suppliers handle stock. *Wrong if:* print-on-demand sellers report equal or higher pain.
- **Behavior:** Most currently reconcile stock by hand in a spreadsheet at least weekly. *Wrong if:* most already use an existing sync app and are happy with it.
- **Willingness to pay:** They'd pay $15-$25/month, against the cost of refunds and bad reviews from overselling. *Wrong if:* most rate an oversell as a minor annoyance with no real cost.
- **Market and channel:** At least 20 qualified sellers can be recruited from two seller communities within two weeks. *Wrong if:* recruiting stalls below 10 qualified respondents.

Every line commits to something specific that could come back false, and each result would send you toward a different next step.

## Common mistakes

- **Bundling.** "Sellers need a better way to manage inventory and would pay for it" holds three hypotheses. Split them.
- **Writing the solution as the problem.** "Users need an AI dashboard" assumes your answer. Describe the problem without your product in it.
- **No kill condition.** If you can't write the result that would prove you wrong, what you have is still a hope.
- **Testing the comfortable ones first.** Solution feedback feels productive, but problem validation is where ideas die, so do that one first.
- **Counting compliments.** "Love this!" tells you the person is polite.
- **Moving the goalposts.** Rewriting the threshold after the data arrives turns research into a search for confirmation.

## Common questions

### What is a startup hypothesis?

A startup hypothesis is a specific, falsifiable statement about your customers, their problem, their behavior, what they'd pay, or your market, written so real customer evidence could prove it wrong. It turns an assumption you're taking for granted into something you can test.

### What's the difference between a startup hypothesis and a scientific hypothesis?

Both are testable claims with a result that could disprove them. They differ in the standard of evidence. A scientific hypothesis usually needs controlled experiments and statistical significance, while a startup hypothesis needs enough real evidence to make a decision, often a clear pattern across 10 to 20 qualified respondents.

### How many hypotheses should a startup test?

Three to seven at the idea stage, one per area that carries real risk. More than ten usually means you're restating the same claim in different words, and a single hypothesis is probably several bundled together.

### Which hypothesis should I test first?

Usually the problem hypothesis. It has the highest impact if it's wrong, and every other hypothesis depends on it. In general, test the one that would collapse the most of your idea if false and that you have the least evidence for.

### How do I know a hypothesis has been invalidated?

When the evidence hits the kill condition you wrote before collecting data. Without a threshold set in advance, you can reread almost any result as "promising."

### Can a hypothesis be partly true?

Yes, and it's common. A problem may be real for one sub-segment and not another, or real but less severe than you assumed. Narrow the hypothesis and test the sharper version instead of counting it as a pass.

### Is Hypothis related to the word "hypothesis"?

Yes. Hypothis is named after the hypothesis, because it runs on hypothesis-driven research. You describe a startup idea, Hypothis turns it into testable hypotheses (problem, willingness to pay, behavior, solution fit, and market), and scores each one against real customer responses.

## Test your hypotheses with real customers

Hypothis turns a raw idea into a falsifiable hypothesis set like the one above. It builds the research round to test it, with screeners and Mom-Test-disciplined questions, and scores each hypothesis confirmed, weakened, or inconclusive from real responses. Those scores roll up into one verdict: Build it, Keep testing, Refine, or Weak demand. It's free during early access. [See how it works](/how-it-works) or [read how to validate a startup idea end to end](/blog/how-to-validate-a-startup-idea).
