From SaaS Idea to Product-Market Fit
You already know how to build software. This guide teaches the other half: how to decide whether something is worth building, and how to know when you were right.
«What evidence should I collect before I spend more time building?»
The loop you will run at every stage. Validation is progressive reduction of uncertainty.
The Fundamental Mistake
A software engineer has an idea for a SaaS product. It feels obviously useful. She spends three months of evenings and weekends building it: clean architecture, tests, billing, a polished landing page. She launches.
Nobody buys it.
Not because the code is bad. The code is excellent. The product works exactly as designed. It fails because of something the code cannot fix:
Building proves that you can build the product. It does not prove that anybody wants it.
Engineers make this mistake for a specific reason: building is the part we are good at. When you are uncertain, you retreat to the activity where you feel competent. Writing code feels like progress. But the risky question was never "can I build this?" — for an experienced engineer the answer is almost always yes. The risky questions are: does anyone have this problem, do they care enough to pay, and can you reach them?
Four words you must keep separate
| Word | What it means | Example |
|---|---|---|
| Problem | Something a specific person struggles with, whether or not your product exists | "Recruiters waste hours coordinating interview times." |
| Solution | One possible way to remove that struggle | Software that automates scheduling |
| Customer | The person who pays to make the problem go away | The agency owner who signs the invoice |
| Market | All the customers with that problem, plus the money that flows around it | Recruitment agencies as a group |
The three-month failure happens when you start at solution and never verify the other three. This guide teaches you to move in the opposite direction: market → customer → problem → evidence → and only then, solution.
Dana is a backend engineer with ten years of experience. One evening she thinks: "Scheduling interviews by email is painful. I could build an AI scheduling tool." Notice: she has started with a solution. We will follow Dana through this entire guide and watch what happens when she works backwards to the market instead.
What Is a Market?
A market is a group of people or companies who share a problem and spend money (or could spend money) on solving it. That's it. Not a report, not a chart — a group of real buyers.
Inside a market, three roles matter. Sometimes one person plays all three. Often they don't.
| Role | Definition | SaaS example |
|---|---|---|
| User | The person who touches the product daily | A recruiter using scheduling software |
| Buyer | The person who approves the payment | The agency owner or operations manager |
| Customer | The account that pays — the company or individual on the invoice | The recruitment agency |
B2B vs B2C
- B2C (business-to-consumer): you sell to individuals. User and buyer are usually the same person. Prices are low, so you need many customers, which means you need cheap distribution.
- B2B (business-to-business): you sell to companies. Prices are higher, customer counts are lower, and buying decisions are more rational: "does this save us money or make us money?" For a solo technical founder, B2B SaaS is usually the more forgiving path — 100 customers paying $200/month is a real business.
Why the user/buyer split matters
In B2B, the person who loves your product often cannot pay for it. A recruiter may adore your tool, but if the agency owner doesn't see the value, no sale happens. The reverse also occurs: a manager buys software the team hates and never uses — which kills you later at renewal time.
Whenever you think about an idea, complete this sentence before anything else: "The user is ___, the buyer is ___, and the buyer pays because ___." If you can't fill in the blanks, that's your first missing piece of evidence.
Dana asks: who would use an interview scheduling tool? Recruiters. Who would pay? Their boss. Why would the boss pay? She realizes she doesn't know. Scheduling is annoying for recruiters, but does it cost the business anything the owner cares about? Unknown. She writes it down as an assumption to test.
What Is a Niche?
A niche is a narrow, specific slice of a market that you deliberately choose to serve first. Narrowing feels like giving up customers. In practice it's the opposite: it's the only way a small founder can win any customers at all.
Watch a market narrow
Each step down does three things:
- Customer discovery gets easier. "Businesses" can't be interviewed. "Shopify fashion brands with high return rates" can be listed by name, found in specific communities, and emailed this week.
- Product design gets easier. A tool for "businesses" must be generic. A tool for fashion returns can ship with the exact workflow, integrations (Shopify, the major return carriers), and vocabulary the customer already uses.
- Marketing writes itself. "Cut return losses for your fashion brand" beats "optimize your business operations" every time.
The tradeoff
- Too broad: you can't name your customers, every prospect needs different features, and you compete with everyone.
- Too narrow: there are only 40 possible customers on Earth, or they have no money, or the problem occurs once a year.
- Useful niche: hundreds-to-thousands of similar customers with a shared painful problem, money to spend, and a place where you can find them.
How to evaluate a niche
Score a candidate niche against these questions. You don't need every answer to be perfect. You do need to notice which answers you don't know.
| Factor | Question to ask |
|---|---|
| Reachability | Could I put my message in front of 100 of these customers within two weeks? Through what channel? |
| Count | Are there at least a few thousand potential customers? (For B2B at $100–500/month, a few thousand is plenty.) |
| Ability to pay | Do these customers have revenue and budgets, or are they broke? |
| Existing software spend | Do they already pay for SaaS? Existing spend proves they buy software. |
| Pain | Do they have problems they complain about, or is everything "fine"? |
| Urgency & frequency | Does the problem hurt this week, and does it recur? |
| Competition | Is there some competition (good — demand exists) but a segment that feels underserved? |
| Trajectory | Is the niche growing, stable, or shrinking? |
| Founder access | Do I know people in this niche, or come from it myself? |
| Founder knowledge | Do I understand their workflow, or can I learn it fast? |
| Distribution channel | Is there a marketplace, community, or channel that concentrates these customers? |
A niche where you can easily reach 500 customers is better than a market containing 5 million customers you cannot reach. Unreachable customers are worth exactly zero to a small founder.
"Companies that schedule interviews" is everyone — too broad. Dana lists narrower options: in-house tech recruiting teams, executive search firms, staffing agencies. She picks staffing agencies that place temporary healthcare workers (nurses, medical techs). Why: her sister runs a desk at one, so she has access; the agencies are mid-sized businesses that already pay for software (applicant tracking systems); and they coordinate enormous numbers of placements, so any coordination pain repeats daily.
Finding Problems
Once you have a niche, resist the urge to invent a solution. Your job now is to hunt for problems that already exist — visible in what people do, not just in what they say.
The vocabulary of pain
- Pain — something that costs real money, time, or stress. "We lose placements because paperwork is late."
- Inconvenience — mildly annoying but tolerable. "The export button is in a weird place." Inconveniences rarely support a paid product.
- Desire — something people want more of (revenue, leads, speed). Desires sell well when tied to money.
- Workflow — the sequence of steps someone follows to get a job done. Problems live inside workflows, so learn the workflow first.
- Workaround — the duct-tape solution people built because no good tool exists. Workarounds are treasure: each one marks a problem someone cared enough to hack around.
- Recurring problem — happens weekly or daily. Recurring pain supports recurring revenue.
- Expensive problem — costs measurable money (lost deals, staff hours, fines).
- Urgent problem — must be solved now, not someday. Urgency is what makes buyers actually move.
Where problems leave footprints
Don't ask "what problems do you have?" — people can't answer that well. Instead, look for these footprints:
- Spreadsheets — every hairy shared spreadsheet is a database and app someone needed but nobody built.
- Manual processes — copy-paste between systems, printed checklists, "then I email it to Carol."
- Agencies and consultants — companies paying outsiders $3,000/month for a service are announcing a budget.
- Virtual assistants and repetitive staff work — a human doing the same steps daily is a process waiting for software.
- Multiple SaaS tools stitched together — Zapier chains and duct-taped integrations mark a gap between products.
- Complaints about existing tools — 1–3 star reviews on G2/Capterra, "does anyone know an alternative to X?" threads.
- Internal tools — when several companies each build the same internal tool, a product is missing from the market.
- Repeated community discussions — the same question appearing monthly in a subreddit or Slack group.
- Feature requests — public roadmaps and forums of competitors show what customers beg for.
- Switching — people migrating between competitors are unhappy people; find out why.
Existing spending is evidence. If someone already spends money or labor solving a problem — a consultant, an employee's hours, a stack of tools, a hand-built script — that is far stronger evidence than someone saying "that sounds useful." Money and hours already flowing are the closest thing you get to proof before you have a product.
Dana visits her sister's agency for an afternoon and just watches. Scheduling interviews turns out to be a five-minute annoyance — the agency uses a shared calendar link and shrugs. But she notices something else: recruiters spend hours every day chasing candidates for compliance documents — nursing licenses, vaccination records, certifications — before a placement can start. There's a giant spreadsheet tracking document status. One agency she later researches pays a part-time admin just to chase paperwork. A nurse who can't start Monday because a document is missing costs the agency the whole placement fee. Spreadsheet + dedicated staff + lost revenue: three footprints on one problem.
Problem Quality: Ranking What You Found
Problem-hunting usually produces several candidates. They are not equal. Rank them with a simple framework — not because the numbers are truth, but because scoring forces you to confront what you don't know.
| Criterion | Question |
|---|---|
| Pain | How much does it hurt? Money lost, hours burned, stress? |
| Frequency | Daily? Weekly? Annually? (Daily beats annually.) |
| Urgency | Must it be solved now, or is it a "someday" problem? |
| Cost | Can you put a dollar figure on the problem? |
| Willingness to pay | Is anyone already paying (tools, staff, consultants) to reduce it? |
| Existing alternatives | Are current solutions absent, bad, or good-and-cheap? |
| Reachability | Can you find the people who have this problem? |
| Market size | Are there enough affected customers to build a business? |
| Founder advantage | Do you have access, knowledge, or skills others lack here? |
Rate each criterion Weak / Unknown / Promising / Strong. Treat "Unknown" as a to-do item, not a zero. The score is a decision aid — a way to compare candidates and expose blind spots — never a mathematical prediction of success.
Three examples
Weak problem
"Developers find it annoying to write commit messages." Real, frequent — but painless, free to ignore, costs nothing, and dozens of free tools exist. Nobody budgets for this.
Interesting problem
"Freelance designers waste time creating proposals." Frequent and mildly costly, and some pay for proposal tools. But the audience is price-sensitive, churns fast, and alternatives are cheap. Possible business, hard fight.
Strong SaaS problem
"Healthcare staffing agencies lose placements because candidate compliance documents are missing or expired." Frequent (every placement), expensive (a lost placement is thousands of dollars), urgent (the shift starts Monday), already paid for (admin staff hours), and the buyers are identifiable businesses with software budgets.
Problem vs Idea vs Product
Engineers blur these four things constantly, and the blur causes bad decisions. Keep them separate:
Why the separation matters:
- You can validate the problem without any solution in mind. (Do agencies really lose placements this way? How often? What does it cost?)
- You can test the solution hypothesis without a product. (If documents were chased automatically, would that actually prevent the losses — or is the real bottleneck the licensing boards?)
- You can be right about the problem and wrong about the solution — or right about both and wrong about the business model. Each layer is a separate assumption, and each gets its own evidence.
When someone says "I validated my idea," they usually mean "I confirmed the problem exists." Those are different layers. Confirming the problem says nothing about whether your specific product, at your price, through your channel, will win. Keep asking: which layer does this evidence support?
What Does "Validate an Idea" Actually Mean?
Validation is not a binary stamp. There is no moment when an idea becomes "validated." There is only evidence — and evidence has grades. Your job is to climb from weak evidence to strong evidence while spending as little time and money as possible.
Behavior beats opinion. What people say is polite, hypothetical, and free. What people do — spend, share, commit, pay, return — costs them something. The more it costs them, the more you can trust it.
The Evidence Ladder
// click any rung to expand it
Everything in the rest of this guide is a technique for climbing this ladder cheaply. Interviews get you to rungs 4–6. Experiments get you to 7–10. A product is only required for rungs 11–13.
Dana maps her evidence so far: rung 5 (the spreadsheet) and rung 6 (the paperwork admin). Encouraging — but it all comes from one agency. Her next assumption to test: this problem is common across healthcare staffing agencies, not unique to my sister's. The cheapest experiment: talk to more agencies.
Customer Interviews
Interviews are the cheapest way to move from rung 2 to rung 6. But most first-time interviews collect garbage, because of one mistake: asking people to predict their future behavior.
- "Would you use a tool that did X?"
- "Would you pay $99/month for this?"
- "Do you think this is a good idea?"
- "How much would you pay?"
People are terrible at predicting their own behavior and excellent at being polite. "Sure, I'd use that" costs them nothing and tells you nothing. This is the core insight of Rob Fitzpatrick's The Mom Test: even your mom would say your idea is great, so ask questions that even a liar can't answer wrongly — questions about facts.
- "Tell me about the last time this happened."
- "What did you do?"
- "Why was that difficult?"
- "How often does it happen?"
- "What happens if you do nothing?"
- "What are you using today?"
- "What does that cost — in money or hours?"
- "Who decides whether to buy software for this?"
These questions have factual answers about the past. Facts can't be polite. If the answer to "when did this last happen?" is "hmm, a few months ago I guess," you just learned the problem is rare — priceless information no prediction question would surface.
A reusable 25-minute interview structure
- Setup (2 min). "I'm researching how <niche> handles <workflow>. I'm not selling anything, and I don't have a product. I just want to understand how it works today." (Not pitching is a superpower: people teach honestly, but they review politely.)
- Map the workflow (8 min). "Walk me through what happens from <start> to <end>." Ask them to be concrete: which tools, which people, which steps. Note every spreadsheet, every copy-paste, every "then I chase them by phone."
- Dig into pain (8 min). Pick the moment they groaned. "Tell me about the last time that went wrong." "What did it cost?" "How often?" "What have you tried?" "What happens if you do nothing?"
- Money & power (4 min). "What do you spend on tools or staff for this today?" "If someone wanted to fix this, who would approve the budget?"
- Close (3 min). "Who else deals with this that I should talk to?" (Referrals compound.) If — and only if — they pulled the conversation toward solutions: "I may test a fix for this. Want to be first to see it?" Their reaction is a small commitment test.
Practical notes: 5 interviews reveal the obvious patterns; 10–15 in one niche is usually enough to rank problems confidently. Take verbatim notes — the customer's exact words become your landing-page copy later. And listen for what they didn't mention: if you have to introduce the problem yourself, that's evidence against it.
Dana emails 30 healthcare staffing agencies ("I'm an engineer researching how agencies handle candidate compliance — 20 minutes, not selling anything") and gets 9 calls. Seven describe recent, specific document disasters without prompting — rung 4. Five maintain tracking spreadsheets — rung 5. Three pay staff specifically for document chasing — rung 6. Interesting twist: nobody complains about interview scheduling. Her original idea is quietly dead, and she's glad she found out in week two instead of month four.
Market Research for Technical Founders
Desk research runs in parallel with interviews. You're not writing a report — you're hunting for specific pieces of evidence. Here's what each source is actually for:
| Source | What evidence to extract |
|---|---|
| Search the exact phrases customers used in interviews. Who advertises on those terms? Ads mean someone profits here. | |
| Reddit & communities | Recurring complaints, workaround threads, "alternative to X?" posts. Frequency of the same complaint over months = persistent pain. |
| YouTube | Tutorial videos on manual processes ("how we track credentials in Excel") prove people struggle enough to seek help. |
| Find your buyers by title; count them; see what they post about. Also your main outreach channel later. | |
| G2 / Capterra | Competitor reviews. Read the 2–3 star reviews first: they list real weaknesses. Note which customer segments complain. |
| App marketplaces | Shopify/Salesforce/etc. app stores show what integrations exist, install counts, and gaps in the ecosystem. |
| Competitor websites | See the inspection checklist below. |
| Job descriptions | Companies hiring a "credentialing coordinator" are paying salaries for your problem. Job posts also name the tools in use. |
| Industry reports | Directional trends only (growing or shrinking niche). Never treat a TAM number as validation. |
| Search trends | Rising or falling interest in the category over years. Useful for trajectory, not for demand proof. |
How to inspect a competitor
"Research competitors" is vague. Inspect these specifics and write them down:
- Customer segment — who is the site talking to? Enterprise? SMB? Which industry words do they use?
- Positioning — what is the one-line promise on the homepage?
- Pricing — plans, price points, per-seat or flat, free tier? This calibrates what the market will pay.
- Reviews & complaints — what do customers love and hate? Hates are your opportunity list.
- Missing functionality — public roadmaps, forum requests, "coming soon" pages.
- Acquisition channels — do they rank on SEO, run ads, sell outbound, live in a marketplace? Their channels hint at what works.
- Integrations — which systems they connect to reveals the tool stack of the niche.
- Switching complaints — "we left X because…" posts are the most honest competitor teardown available.
Dana finds two big credentialing platforms — both enterprise-priced, built for hospitals, with G2 reviews from staffing agencies saying "overkill for us," "sales process took months," "we just went back to spreadsheets." She also finds 40+ job posts for "credentialing coordinators" at staffing agencies. Translation: demand is proven, and the small-agency segment is underserved. That's a wedge.
Competition
Beginner instinct: "someone already built it, so I shouldn't." That instinct is wrong often enough to be dangerous. Competition is evidence that customers pay for solutions to this problem. A market with zero competitors demands an uncomfortable question: is it undiscovered — or is there no money in it? Usually it's the second.
| Situation | What it usually means | Your move |
|---|---|---|
| No competitors, no substitutes, no spending | Most likely: nobody wants this solved badly enough to pay | Demand extraordinary evidence before proceeding |
| No software competitors, but heavy manual spending (staff, consultants, spreadsheets) | Genuine gap — the money exists but flows to labor | Strong opportunity; move fast on validation |
| Crowded market, similar products | Demand proven, differentiation hard | Only enter with a specific underserved segment or sharp angle |
| Big generic players, unhappy niche segments | The classic indie SaaS wedge | Serve one segment 10× better than the generalists do |
| Competitors visibly growing and hiring | Market is expanding — rising tide | Healthy sign, if you have a differentiated position |
Worked examples
- Basecamp existed; project management looked "done." Then Trello won with boards, Asana with teams, Linear with software engineers specifically. Crowded market, repeated wins through segment + angle.
- Mailchimp dominated email; ConvertKit still built a business by serving one segment — creators — better.
- Countless "Uber for X" ideas had zero competitors. Most died: no competitors because no demand.
Be afraid when: the incumbents are good and cheap and loved (read their reviews — is the hate real?); switching costs lock customers in (deep data or workflow lock-in); a free tool from a platform vendor (Google, Microsoft, Shopify) covers 80% of the job; or your only differentiator is "mine will be better built." Customers don't buy architecture.
Distribution: Can You Reach Anyone?
Engineers systematically underestimate distribution. Here is the uncomfortable truth: if you cannot reach potential customers, you cannot validate the idea — and you certainly cannot sell the product. A brilliant product with no path to customers is a hobby.
That's why distribution appears here, before the chapters on building anything. Reachability is a validation question, not a marketing afterthought.
«Where can I find 20 people with this problem this week?» is more useful to an early founder than «How large is the TAM?». TAM (total addressable market — the theoretical revenue if everyone bought) describes a market you don't have. The 20-people question tests the channel you actually need: first for interviews, then for pilots, then for sales. If you can't answer it, that is your riskiest assumption right now.
Founder–channel fit
Channels are not interchangeable. Each rewards different skills, and you should pick the one that fits you and your niche:
| Channel | Works when… | Fits founders who… |
|---|---|---|
| Outbound email | B2B, identifiable buyers, clear ROI story | Can write concisely and tolerate rejection |
| Buyers live there (most B2B roles do) | Are willing to post and DM consistently | |
| Communities | The niche gathers in Slack/Discord/forums/associations | Genuinely participate rather than spam |
| SEO / content | Customers search for the problem; you can wait 6–12 months | Like writing; play long games |
| Marketplaces (Shopify, Salesforce, Chrome…) | Your product extends a platform your niche already uses | Build integrations well — a real engineer advantage |
| Partnerships / agencies | Someone else already sells services to your niche | Can offer partners a margin or a service upgrade |
| Existing audience | You already have followers/newsletter in the niche | Have been building in public |
| Integrations | Being listed in another tool's directory brings warm traffic | Ship reliable APIs fast |
| Paid ads | You know your numbers (price, conversion, churn) — rarely true pre-PMF | Have budget to burn on learning |
For a solo technical founder in B2B, the honest default is: outbound email + LinkedIn + one community, later reinforced by a marketplace or integration listing. Pick one primary channel and get good at it before adding another.
Dana's channel test: can she find 20 healthcare staffing agencies this week? She builds a list of 120 from a state registry and LinkedIn in one evening. Her cold-email reply rate for research calls was 30% (helped by mentioning her sister's agency — founder access at work). Channel: proven. She notes that agencies also cluster at two industry associations — a future partnership route.
Validation Experiments, Cheapest First
You now have a ranked problem, a defined customer, and a channel. Time to test the solution side. The tools below are ordered by cost. At each step ask: what is my riskiest remaining assumption, and what is the cheapest experiment that could kill it? Then run that experiment — not a more expensive one.
1 · Desk research — no product
Tests: "This problem exists beyond my head." Produces: rung 2 evidence — complaints, workarounds, competitor demand, job posts. Cannot prove: that anyone will change behavior or pay. Use: always, first, for a few days at most.
2 · Customer interviews — no product
Tests: "The problem is painful, frequent, and costly for this specific niche." Produces: rungs 3–6 — stories, workarounds, existing spend, plus the vocabulary and workflow you'll need later. Cannot prove: that your solution is right or that anyone will pay you. Use: always, before any building. 10–15 interviews per niche.
3 · Landing page — a promise only
One page describing the problem and promised solution, with a call to action (join waitlist / book a demo). Tests: "The message resonates enough that strangers act." Produces: weak-to-medium interest signals; email signups sit around rung 3. Cannot prove: willingness to pay — an email address costs nothing. Use: to test messaging and collect interview candidates; never as proof of demand by itself.
4 · Fake door / smoke test — measure a click
A "Buy now — $99/mo" or "Start trial" button that leads to "We're onboarding soon, leave your email." Or a feature button inside an existing product that measures clicks before the feature exists. Tests: "People take a purchase-shaped action, not just an interest-shaped one." Produces: stronger intent signal than a waitlist. Cannot prove: actual payment or retention. Use: when traffic exists and you want to compare offers/prices cheaply. Be honest with people immediately after the click.
5 · Demo or clickable prototype — show the workflow
Figma clickthrough, slides, or a hard-coded UI with fake data. Tests: "When customers see the actual workflow, do they lean in — and does it match how they really work?" Produces: rung 7 (agreement to test) and priceless workflow corrections. Cannot prove: real usage; demos flatter every product. Use: in sales conversations after interviews confirmed the problem.
6 · Concierge MVP — you are the product
Deliver the outcome completely manually and visibly: you personally chase the documents, build the report, do the matching. No software. Tests: "Is the outcome valuable enough that customers want it (ideally pay for it)?" Produces: rungs 8–10 evidence plus a perfect specification — after doing the job by hand 20 times, you know exactly what to automate. Cannot prove: that customers will accept a self-serve tool, or that unit economics work at scale. Use: B2B services-like workflows; the highest learning-per-hour experiment available to an engineer.
7 · Wizard-of-Oz MVP — looks automated, isn't
Customer sees a normal product; behind the curtain, humans do some steps. Tests: "Do customers engage with the product-shaped experience repeatedly?" Produces: real usage data (rung 11 signals) without building the hard parts. Cannot prove: feasibility or cost of the real automation. Use: when the risky assumption is demand, not technology — and be sure the manual work is actually automatable later.
8 · Paid pilot — money for an early version
A customer pays (even a modest amount) for a defined trial period with agreed success criteria. Tests: "Will a real budget-holder exchange money for this outcome?" Produces: rung 10 — the strongest pre-product evidence there is. Cannot prove: retention or repeatability across many customers. Use: B2B, once a demo or concierge run has landed. Charging for pilots also filters out tire-kickers.
9 · Pre-sale — money before the product is complete
Customers commit cash (deposit, discounted annual plan) before the finished product exists. Tests: "Is the pain urgent enough to pay now and wait?" Produces: rung 10 across multiple customers; funds development. Cannot prove: that you can deliver, or that usage will follow. Use: when several pilots or demos have generated pull. Refund anyone if you can't deliver.
10 · MVP — the smallest real product
Covered fully in Level 14. Tests: "Will customers use this repeatedly, in real life, without me pushing?" Produces: rungs 11–12 — usage and retention, the evidence no cheaper method can generate. Cannot prove: PMF by itself; that requires observing retention over months. Use: only when the remaining risky assumptions genuinely require working software.
Each method is only meaningful after the cheaper ones. A paid pilot with the wrong niche wastes months; a landing page before interviews tests words you made up. Climb in order, skip only with a reason.
Dana's riskiest assumption isn't the problem anymore (verified) — it's "agencies will trust an outsider's tool with compliance." Cheapest test: a concierge run. She offers three agencies: "For four weeks I'll chase and track all candidate documents for you — flat $400." Two say yes. One month in, a failed assumption surfaces: her pitch was 'never lose track of documents,' but what makes the owners lean forward is expiry tracking — licenses silently expiring on already-placed workers, which creates compliance risk during audits. She narrows the concept: not generic document collection, but credential expiry monitoring with automated candidate chasing, for healthcare staffing agencies.
When Should I Start Coding?
The question every engineer actually cares about. The answer is not "never build before validation" — that advice overcorrects and ignores that for you, building is cheap. The real principle:
Build only when building is the cheapest way to test the next important assumption. Sometimes a two-day prototype is exactly that — the cheapest experiment. A three-month product almost never is.
What you ideally know before serious building
- Who has the problem — a customer you can describe in one sentence
- What the problem is — in the customer's own words, without mentioning your solution
- How they solve it now — the workaround, tool, or staff currently doing the job
- How painful it is — a cost in money or hours, from real examples
- How often it happens — frequency supports subscription revenue
- Who pays — the budget-holder, and why they'd approve it
- How you reach them — a channel you have personally tested
- Why existing solutions are insufficient — from reviews and interviews, not your guess
- Meaningful commitment — at least one person on rung 8+ (time, data, LOI, or money)
You will rarely have all nine. Treat missing ones as labeled risks you're consciously accepting — not as details to discover after launch.
Legitimate reasons to code early
- The riskiest assumption is technical feasibility ("can an LLM actually extract expiry dates from scanned licenses at 99% accuracy?"). Then a throwaway spike is validation.
- A rough working demo will unlock conversations that slides can't — common when selling to skeptical technical buyers.
- The prototype takes days, not weeks, and you commit to throwing it away if evidence says so.
"I'll add SSO, multi-tenancy, and a billing system first, then start talking to customers." Translation: coding feels safe, and talking to strangers doesn't. The codebase grows; the evidence doesn't.
"Agencies won't trust OCR on licenses. I'll spend three days testing extraction accuracy on 50 sample documents. If it's under 95%, this product changes shape." A dated, falsifiable, timeboxed test.
MVP: The Minimum Viable Product
An MVP is not "version 1 with fewer features." It is:
The smallest real solution capable of testing the most important remaining assumption — usually: "will customers use this repeatedly and pay?" (This is Eric Ries's original meaning: minimum viable learning, delivered through a product.)
Engineers overbuild MVPs in predictable ways: infrastructure for scale that hasn't arrived, settings pages nobody asked for, edge-case handling for users who don't exist, three integrations when the niche uses one, and a rewrite in a "better" framework halfway through. Every one of these delays the only thing that matters: contact with reality.
Bad MVP vs good MVP — same idea
Idea: credential expiry tracking for healthcare staffing agencies.
- Multi-tenant architecture, SSO, role-based permissions
- Integrations with 6 applicant-tracking systems
- ML pipeline for automatic document classification
- Configurable workflow builder
- Native mobile app
- Self-serve signup and Stripe billing
Tests nothing new. Every feature answers "can I build this?" — a question that was never in doubt.
- Upload/import a candidate list (CSV beats an integration)
- Track each credential with an expiry date
- Automated email/SMS chasing of candidates for documents
- One dashboard: "who is expiring in the next 30 days"
- Manual onboarding — Dana sets up each agency herself
- Invoices sent by hand
Tests exactly one thing: do agencies use this weekly and pay for it? Everything else waits for evidence.
Two rules of thumb: your MVP should embarrass you a little (Reid Hoffman's line — if it doesn't, you shipped late), and it should be usable for the core job without you in the room by roughly week two of a customer's life. "Minimum" is not an excuse for broken — it means narrow, not shoddy: one workflow, done reliably.
First Customers
With an MVP live, the game changes from validating to selling — and the two feel different. Validation conversations ask "is this real?"; sales conversations ask "will you buy this, this month?" Many founders stall here, endlessly "validating" to avoid the discomfort of asking for money. Once you have rung-8+ evidence, start asking.
5–20 strong early customers beat thousands of free signups. A strong early customer matches your target niche, feels the pain now, pays something, uses the product, and talks to you. Free signups from the wrong audience generate noise, support load, and misleading feature requests — negative value at this stage.
What this stage looks like in practice
- Founder-led sales. You, personally, emailing and demoing. Nobody can sell an unproven product but the founder, and every objection you hear is product research.
- Manual onboarding. Set up each customer yourself, on a call. You'll watch where they get confused — that's your real bug list.
- Watching customers use the product. Screen-share sessions beat analytics at this scale. Where do they hesitate? What do they ignore?
- Support as research. Answer everything personally. Support tickets are interviews that customers initiate.
- Usage observation. Instrument the 3–5 events that define "the product did its job," and check them per-customer weekly, by name — not as aggregate charts.
All of this is deliberately unscalable, and that's fine. As Paul Graham argues in "Do Things That Don't Scale," early startups are hand-cranked; the recruiting of the first users is almost always manual. You're not building a machine yet — you're learning what machine to build.
Dana converts her two concierge agencies into paying software customers at $250/month, then works her interview list and association contacts. Six weeks of founder-led sales later: 7 paying agencies. She onboards each on a video call and keeps a spreadsheet: per agency, documents chased per week and dashboard logins per week. Two agencies barely log in — she calls them to find out why instead of guessing.
Idea Validation Is NOT Product-Market Fit
This is the distinction founders most often blur, so let's make it sharp.
Evidence that:
- the problem exists
- customers care about it
- they may adopt or pay for a solution
It's evidence about a problem and an intention. It can all be true and the product can still fail.
Evidence that:
- a specific product
- for a specific market
- repeatedly delivers enough value
- that customers keep using and keep buying it
It's evidence about sustained behavior with a real product. Marc Andreessen's original description: being in a good market with a product that can satisfy that market — and you can feel when it's happening because demand starts pulling.
Between validation and PMF sit several stages, each requiring its own evidence:
Why this matters practically: pre-sales, LOIs, and waitlists live in "demand validation." They justify building. They do not justify scaling, hiring, or spending on growth — decisions that should wait for retention evidence. Skipping stages is how funded startups die with beautiful demand metrics and empty week-8 dashboards.
Measuring Product-Market Fit
PMF isn't a single number, but it leaves measurable fingerprints. Watch these:
- Activation — the share of new customers who reach the product's first moment of real value (e.g., first automated document chase sent). Low activation means the product promises value it doesn't deliver, or delivers it too slowly.
- Repeated use — do customers return at the natural frequency of the problem? A daily-pain product should see weekly-or-better use.
- Retention — the share still active after 1, 3, 6 months. The core PMF metric.
- Churn — the share who cancel each month. Early-stage B2B churn above ~5%/month is a flashing warning; low single digits or better is healthy.
- Expansion — existing customers adding seats, usage, or upgrades. Expansion means the product spreads inside accounts.
- Organic referrals — new customers who arrive saying "X told me about you," without you asking.
- Willingness to pay — price resistance dropping; customers accepting renewals and increases without drama.
- Customer pull — inbound demos, feature demands phrased as "when can we have," people chasing you.
- Sales getting easier — shorter cycles, higher close rates, less convincing.
- Complaints during downtime — when your product breaks and customers are angry within minutes, congratulations: you're load-bearing.
Cohort retention, simply
Group customers by start month (each group is a cohort). For each cohort, plot the % still active (or still paying) after month 1, 2, 3… The shape tells you everything:
100% → 60% → 35% → 18% → 9% → …
The curve keeps sliding. Whatever growth you buy leaks out. No PMF — fix the product or the market before adding customers.
100% → 70% → 62% → 60% → 59% → 59%
The curve goes flat: a durable core of customers has made the product a habit. A flattening retention curve is the closest thing to a PMF signature that exists.
The Sean Ellis question
Ask active users: «How would you feel if you could no longer use this product?» with options Very disappointed / Somewhat disappointed / Not disappointed. Sean Ellis, after studying many startups, observed that companies where ≥40% answered "very disappointed" generally went on to grow well, and those far below struggled. Rahul Vohra's team at Superhuman famously turned this into an operating system: measure the score, segment for the users who love you, and split the roadmap between doubling down on what they love and fixing what blocks the fence-sitters — moving their score from 22% to 58% in about a year.
40% is a benchmark from observed correlations, not a universal law. It's sensitive to who you survey and how many. Use it as a thermometer and a segmentation tool. When survey scores and behavior disagree, trust behavior: retention and renewals are customers voting with money and time; surveys are customers being asked to have feelings on demand.
Six months in, Dana has 19 agencies. Her monthly cohorts flatten around 80% by month 3. Two of the early churns shared a trait: tiny agencies (<20 placements/year) where the pain was too rare — sharpening her ideal customer profile rather than scaring her. Her Sean Ellis survey: 9 of 17 respondents "very disappointed" (53%). Three new agencies arrived by referral. Signals are aligning.
What PMF Feels Like
- You push the product uphill; every sale is a fight
- People sign up, then disappear
- Feature requests point in twenty directions
- You keep rewriting the homepage because the target customer is unclear
- Retention curves slide toward zero
- Constant repositioning; "maybe we should also…"
- A clear customer segment with a clear use case
- Repeat usage without prompting
- Retention curves flatten
- Customers actively ask for the product and for more of it
- Referrals appear on their own
- Customers complain loudly when the product fails or capacity runs out
- Demand starts pulling the company; the bottleneck shifts from selling to delivering
Three honest caveats, because PMF is romanticized:
- It's rarely a single magical moment. More often it's a gradual shift: sales calls stop feeling like begging, churn quiets down, the inbox changes tone. You notice it in hindsight.
- PMF has strength. Weak PMF (a retention floor at 40%, lukewarm pull) and strong PMF (customers evangelizing, waitlists) are different worlds. You can operate at weak PMF while working to deepen it.
- PMF can be lost. Markets shift, competitors improve, platforms change rules. PMF is a state you maintain, not a trophy you keep.
Continue, Narrow, Pivot, or Kill
Every experiment ends in a decision. There are exactly four, and choosing consciously — on a schedule, against pre-written criteria — is what separates systematic founders from drifters.
| Decision | When the evidence says | Example |
|---|---|---|
| Continue | Assumptions are holding: pain confirmed, commitments appearing, retention forming. Move to the next assumption on the list. | Pilot customers renew → build the MVP. |
| Narrow | The problem is real, but the strongest demand comes from a sub-segment. Most "meh" results hide a "wow" segment. | Fashion brands overall are lukewarm, but high-return-rate brands over $1M are desperate → serve only them. |
| Pivot | A core assumption failed, but the work surfaced a different strong problem in the same territory. | Dana: scheduling was a dead end; document chasing was the real fire. Same niche, new problem. |
| Kill | Repeated experiments show weak demand: no urgency, no spending, no commitments — across multiple honest attempts. | Fifteen interviews, two channels, zero rung-8 evidence → stop. |
Killing a weak idea early is a successful validation outcome. The purpose of validation is not to bless your idea — it's to find the truth cheaply. Discovering "no" for $0 and three weeks, instead of $0 revenue and six months of nights and weekends, is the system working perfectly. Engineers understand this instinctively in another domain: a failing test that catches a bug before production is a win.
Practical guardrails: decide in advance what result would change your mind ("if fewer than 2 of 10 interviews reveal existing spend, I stop"); timebox each stage; and beware the two failure modes — quitting after one awkward interview, and "just one more feature" for a year. Evidence, not mood, makes the call.
The Complete Validation System
Everything above compresses into one reusable pipeline. This is the mental model to keep:
At every arrow, the same loop runs: Assumption → Experiment → Evidence → Decision (continue / narrow / pivot / kill). The pipeline is linear on paper and loopy in real life — you'll revisit stages. That's normal. What must never happen is skipping from stage 8 to stage 12 on vibes.
Case Study: Dana and ShiftDocs
Dana's story has run through every level. Here it is end to end — roughly nine months, one engineer, no funding.
| Stage | What happened | Evidence rung |
|---|---|---|
| Random idea | "AI interview scheduling" — a solution looking for a problem | 1 |
| Broad market | "Companies that hire" — unusable; everyone and no one | — |
| Niche | Healthcare staffing agencies: founder access (sister), existing software spend, daily coordination volume | — |
| Customer research | An afternoon watching a real agency work | 2→4 |
| Problem discovery | Scheduling: shrug. Compliance document chasing: spreadsheets, a paid admin, lost placements | 5–6 |
| Interviews | 9 agencies; 7 unprompted disaster stories; 3 with dedicated paperwork staff | 4–6 |
| Competitor research | Enterprise credentialing platforms exist; small agencies call them "overkill" and revert to spreadsheets | demand proven, segment underserved |
| Validation experiment | Concierge: "$400 flat, I chase your documents for 4 weeks" — 2 of 3 agencies accept | 10 |
| Failed assumption | Pitch was "organized document collection"; owners actually light up at expiry risk — licenses silently lapsing on placed workers, an audit nightmare | learning |
| Narrowed idea | ShiftDocs: credential expiry monitoring + automated candidate chasing, for healthcare staffing agencies | — |
| Paid pilot | Both concierge agencies agree to pay $250/mo for a software version, with success criteria: "zero placements delayed by missing documents this quarter" | 9–10 |
| MVP | 3 weeks: CSV import, expiry tracking, automated email/SMS chasing, one "expiring in 30 days" dashboard. Manual onboarding, manual invoices | — |
| First customers | Founder-led sales through the interview list and industry association → 7 paying, then 19 by month six | 10–11 |
| Retention | Cohorts flatten near 80% at month 3; churned accounts are all tiny agencies → ICP sharpened to 50+ placements/year | 12 |
| PMF signals | 53% "very disappointed"; 3 referral customers; angry calls within an hour of an outage; feature requests converge on one theme (audit reports) | 12–13 |
Note what Dana never did: she never built the scheduling tool, never bought ads, never wrote a business plan, and never spent more than three weeks building before the next contact with reality. Her only unusual skill was willingness to run the loop: assumption → experiment → evidence → decision, about fifteen times.
A Failure Story: Validation Theater
Now watch Marco, a talented full-stack developer, do every step in a way that feels like validation but isn't.
| What Marco did | What he concluded | What he misread |
|---|---|---|
| 1 · Thinks of an AI SaaS idea: "AI meeting summarizer for teams" | "This is obviously useful — I'd use it" | Rung 1 evidence. He started from a solution, and "I'd use it" describes one unusual user: him. |
| 2 · Reads a market report: "meeting productivity software = $4B by 2027" | "Huge market → huge opportunity" | TAM measures a market, not his access to it. Also: a hot AI category means maximum competition, including free features from Zoom, Teams, and Meet themselves. |
| 3 · Asks friends and coworkers whether they like the idea | "Everyone said it's cool" | Prediction questions to people who like him. Politeness data. Nobody was asked about past behavior, current spend, or budgets. |
| 4 · Ships a landing page with AI-generated copy | "Ready to measure demand" | The copy came from his head, not from customer interviews — so the page tests words no customer ever said, aimed at no specific niche. |
| 5 · Gets 300 email signups from Product Hunt and Hacker News | "300 signups = demand!" | Free signups from tech-curious browsers, not from any defined buyer. Rung 3 at best. No one was asked for time, data, or money. |
| 6 · Declares the idea validated | "Green light" | He has zero evidence above rung 3, no niche, no buyer, no price signal, no channel — and doesn't know it. |
| 7 · Spends four months building: transcription pipeline, speaker diarization, Slack + Notion + calendar integrations, team billing | "Making great progress" | Progress on the only risk he never had: whether he could build it. Meanwhile the market moved and platform vendors shipped native summaries. |
| 8 · Launches to the waitlist. 41 people activate; almost nobody returns; nobody pays | "Maybe I need more features?" | The retention curve — straight to zero — is the market answering clearly. More features can't fix "low-priority problem, free alternatives." |
| 9 · Post-mortem interviews (the ones he skipped) reveal: mildly annoying problem, already handled "well enough" by built-in tools, no budget owner anywhere | "The problem was low priority" | Every fact he learned in month six was learnable in week one, for the price of ten conversations. |
Marco never lied to himself on purpose. He collected real numbers (a $4B report, 300 signups) — they were just numbers that measured the wrong thing. Validation theater is doing activities that resemble validation while carefully avoiding the experiments that could say no: asking about past behavior, asking a specific niche, and asking for commitment.
SaaS Idea Validation Scorecard
Use this every time you evaluate an idea. Click a rating for each factor; be brutally honest, and remember that Unknown is the most useful rating — it's your research to-do list. The score is a decision aid, not a prediction.
Before I Spend a Month Building This
Answer these in writing. Unanswered questions aren't blockers — they're your next experiments.
- Can I describe the customer in one sentence? "Healthcare staffing agencies placing 50+ workers/year" — not "businesses."
- Can I describe the problem without mentioning my solution? If the problem only exists as "they lack my product," it isn't a problem.
- Have I seen evidence this problem exists independent of me? Complaints, reviews, job posts, communities.
- Have I spoken to at least 10 people who experienced it recently?
- Did they tell me specific, recent stories without prompting?
- What do they do about it today? Name the tool, spreadsheet, staff member, or "nothing."
- What does the current approach cost them, in money or hours?
- How often does the problem occur? Daily/weekly supports subscriptions; yearly rarely does.
- What happens to them if they do nothing? If the honest answer is "not much," urgency is missing.
- Who is the user, and who is the buyer? Are they different people?
- Who controls the budget, and roughly how big is it?
- Why are existing solutions insufficient — in customers' words, not mine?
- Can I list 20 potential customers by name today?
- Which channel reaches them, and have I personally tested it?
- Have I asked anyone for a meaningful commitment? Time, data, an LOI, a pilot, money.
- Did anyone give one? Rung 8+ on the Evidence Ladder.
- What price am I hypothesizing, and what existing spend anchors it?
- Will this product deliver value repeatedly, month after month? Retention potential — or is it a one-and-done tool?
- What is the single riskiest assumption right now?
- What is the cheapest experiment that could prove it false?
- Is coding actually required to run that experiment? If no: don't code yet. If yes: build the smallest thing that runs it.
- What result would make me narrow, pivot, or kill — decided in advance?
- What did the last experiment teach me, and what decision did I take from it?
The Software Engineer's SaaS Validation Playbook
The whole guide, condensed. Eight stages: Find → Verify → Test → Sell → Build → Observe → Retain → Scale. Print this section; it's your permanent reference.
Misconceptions, Corrected
"I would use it, therefore other people will."
You are a sample of one, with rare skills, rare tolerance for rough tools, and total bias. Your enthusiasm is a hypothesis (rung 1), never evidence. Ask instead: who else has this problem, and what do they currently do about it?
"There are no competitors, so the opportunity is huge."
Usually there are no competitors because there's no money. Empty markets must prove themselves with heavy manual spending (staff, consultants, workarounds). No competitors and no substitutes and no spending = a warning, not a jackpot.
"There are competitors, so the opportunity is gone."
Competitors prove customers pay. The question is whether a segment is underserved — generic incumbents almost always leave niches behind (see Level 10). Read their 2-star reviews and find out who they're failing.
"People said they like the idea, therefore it is validated."
Liking is free. Validation is behavior: workarounds, spend, time, data, money (rungs 5+). If your best evidence is compliments, you have no evidence yet.
"1000 waitlist signups means PMF."
A signup is a rung-3 signal about curiosity. PMF is retention of a real product months in. A waitlist can justify building an MVP; it can never substitute for watching whether customers stay.
"A large TAM means a good SaaS opportunity."
TAM describes a market's theoretical size, not your access to it. A $10B market you can't reach is worth less to you than 500 reachable customers with a burning problem. Ask "where do I find 20 this week?" first.
"I need an MVP before talking to customers."
Backwards. Interviews need no product — and they're what tells you which MVP to build. Talking first is also cheaper by two orders of magnitude. Not having a product makes interviews better: you learn instead of pitching.
"MVP means a badly built product."
MVP means narrow, not broken: one core workflow, done reliably, everything else omitted. A buggy product tests nothing except customers' patience. Minimum refers to scope; viable refers to quality of the one job it does.
"More features will create PMF."
If the core value doesn't retain customers, extra features widen a product nobody needs. Feature requests pointing in twenty directions is a symptom of a missing core, not a roadmap. Fix retention of the core job first (Superhuman's method: double down on what lovers love).
"If the product is good, customers will find it."
They won't. Nobody is searching for your product; at best they search for the problem. Distribution is half the business, and if you can't reach customers you can't even validate. Pick and test a channel before building (Level 11).
"Product-Market Fit can be validated before a product exists."
By definition, no. You can validate the problem, the solution direction, and demand pre-product — and you should. But PMF is a property of a specific product repeatedly delivering value in a market, evidenced by retention. No product, no PMF (Level 16).
"Revenue automatically means PMF."
Founder-led hustle can generate revenue from a leaky product for a while. If the retention curve slides toward zero, that revenue is a treadmill, not fit. Watch renewals and cohort curves, not just MRR.
"A clever AI feature is a business."
A feature is a business only when attached to a customer, a painful problem, a price, and a channel. Clever capabilities without those get absorbed into platforms (ask Marco). Start from the problem; let AI be the how, not the what.
"Technical difficulty creates customer value."
Customers pay for outcomes — time saved, money earned, risk removed — and are entirely indifferent to how hard it was to build. Some of the best SaaS businesses are technically boring. Difficulty can be a moat, but only after value exists.
Sources / Further Reading
The ideas here stand on well-known shoulders. Where the frameworks disagree, it's mostly about emphasis: Blank and Fitzpatrick push discovery-before-building hardest; Ries emphasizes build-measure-learn loops with minimal products; Andreessen emphasizes market quality over everything; Ellis and Vohra focus on measuring fit once a product exists. This guide sequences them: Blank/Fitzpatrick first, Ries in the middle, Ellis/Vohra at the end.
- Marc Andreessen — "The Only Thing That Matters" (the origin of PMF as a concept): pmarchive.com/guide_to_startups_part4.html
- Rahul Vohra — "How Superhuman Built an Engine to Find Product/Market Fit" (First Round Review): review.firstround.com
- Sean Ellis — the "very disappointed" survey and the 40% heuristic; see Hacking Growth and pmfsurvey.com
- Rob Fitzpatrick — The Mom Test (how to ask questions that produce facts): momtestbook.com
- Steve Blank — The Four Steps to the Epiphany and customer development: steveblank.com
- Eric Ries — The Lean Startup (MVPs, build-measure-learn, validated learning)
- Paul Graham — "How to Get Startup Ideas": paulgraham.com/startupideas.html and "Do Things That Don't Scale": paulgraham.com/ds.html
- Y Combinator — Startup School (free course; see especially the lectures on talking to users and finding product-market fit): startupschool.org
- First Round Review — deep founder case studies: review.firstround.com