A prospect says the idea is brilliant. Another clicks the announcement, joins the waitlist, and forwards it to a colleague. A third tries the workflow once but never returns. These are all signals of interest. They are not interchangeable evidence, because each one supports a different decision and leaves different explanations alive.
What actually counts depends on the claim being made. Attention can justify testing a promise more closely. Intent can justify arranging a real trial. Activation can justify improving the path to first value. Return can justify investing in a repeatable workflow. Payment can justify a commercial test. No single signal proves the whole product, and even a strong signal becomes weak when it is used to support a decision it never touched.
Compliments and clicks are clues, not verdicts
Compliments are socially easy. They can reveal language that resonates, an outcome people admire, or a relationship worth continuing. They can also reflect politeness, support for the founder, enthusiasm for AI, or relief that someone understands the problem. A compliment justifies asking a sharper question. It does not justify projecting usage or revenue.
Clicks have a clearer behavioral component, but the behavior is still small. A person noticed a promise and spent a moment investigating it. That can help compare messages within the same audience and channel. It does not show that the person has the problem, can adopt the workflow, will trust the result, or will pay. The click is honest evidence of attention when it is allowed to remain evidence of attention.
The common error is not collecting weak signals. Cheap signals are useful because they help direct expensive tests. The error is promoting them after the fact. A waitlist gathered from a broad launch may be a good recruiting pool for interviews. Calling it validated demand removes the very questions those interviews need to answer.
Five signal families answer five different questions
Attention asks whether the right people notice and understand the promise. Useful observations include qualified visits, replies from the intended role, or time spent examining a concrete use case. The honest decision is whether to keep investigating that audience-message pair. Attention cannot justify building the workflow behind the promise.
Intent asks whether interest survives a meaningful next step. Booking a working session, completing a detailed intake, sharing a permitted sample, inviting a stakeholder, or accepting pilot terms all consume something scarce. These behaviors can justify a real delivery test. They still cannot show that the proposed outcome works or matters after use.
Activation asks whether a person reaches the first outcome the product promises. For an AI research workflow, opening the application is not activation; producing a decision-ready brief from a live question might be. Activation can justify improving delivery, onboarding, and reliability around that moment. A single successful outcome does not establish a habit.
Return asks whether the value is strong and recurring enough for someone to come back without being carried by the test team. The return should match the natural rhythm of the job. Daily repetition makes no sense for quarterly planning, while a second session arranged only through persistent reminders says little about pull. Credible return evidence can justify investment in a repeatable product loop.
Payment asks whether the value can cross a commercial boundary. It is strong evidence for price, buyer, and purchasing motion when the money comes from the intended customer under realistic terms. It may be weak evidence for retention or scalable delivery. A paid bespoke project can support a service business while saying very little about a self-serve product.
Evidence strength rises when alternatives disappear
A signal is strong when fewer plausible explanations can produce it. A page click might come from curiosity, confusing copy, novelty, or genuine need. A completed live workflow after a user provides sensitive input has fewer explanations. A voluntary return for the next case narrows them further. Strength comes from the connection between behavior and claim, not from how impressive the dashboard looks.
Stronger tests usually cost more. They require realistic inputs, careful delivery, time from the intended user, and sometimes money changing hands. That does not make the most expensive test automatically best. The aim is to buy enough evidence for the next decision at the lowest responsible cost. Testing retention before anyone can activate wastes effort; treating cheap attention as retention wastes judgment.
Match test cost to decision cost. Rewriting a landing-page promise is cheap and reversible, so attention may be sufficient. Building a workflow that handles confidential records is expensive and harder to unwind, so qualified intent and real-input activation should come first. Expanding a team around a repeat-use product calls for evidence that people return in conditions resembling normal use.
False positives hide inside the test design
The wrong audience is a frequent source. Builders, investors, friends, and early adopters may engage because they enjoy new technology. Their attention can be sincere while remaining unrelated to the intended buyer’s problem. Segment results by role, context, and problem history before combining them into one encouraging total.
Founder assistance creates another false positive. White-glove support is useful when it is the experiment, but it can conceal onboarding effort, missing context, and weak motivation. Record what the participant completed alone, what required prompting, and which value came from the founder rather than the proposed product. Concierge delivery should expose the workflow, not quietly manufacture enthusiasm.
Incentives can distort the signal as well. A participant paid for research may complete a long task because the research payment is valuable. A free pilot may attract a team that would never enter procurement. An unusually generous setup can turn a weak product into a pleasant service. Incentives are not forbidden; they simply become part of the explanation and limit what the result can justify.
Novelty is especially potent in AI products. A surprising output can earn screenshots and shares before it earns a place in recurring work. To separate novelty from value, repeat the same job with a less exciting but realistic input, wait for the job’s natural recurrence, and observe whether the user initiates the next session. The second use is often less photogenic and more informative.
What each signal can honestly justify
Consider an illustrative compliance concept. An operations audience clicks a page promising faster policy comparison. That supports further investigation of the promise. Several qualified managers then upload permitted sample documents and schedule a review session. That supports testing delivery with real inputs. Neither result supports a claim that the output will enter an accountable compliance decision.
Now suppose reviewers use the assisted comparison on live cases and reach an accepted first result with corrections they can explain. Activation supports investing in source visibility, correction tools, and more varied cases. If the same reviewers return when new policies arrive, without the founder arranging every session, the team has evidence for a recurring workflow. It still needs to discover who buys and under what terms.
If one department pays for a narrowly defined pilot, the team has evidence that this buyer can cross a purchasing boundary for this outcome. The pilot does not prove that another department will buy, that gross margins will work, or that users will remain after the founder steps away. The conclusion can be both encouraging and narrow. Honest evidence often sounds less dramatic because it preserves the next uncertainty.
Turn signals into a decision record
For each experiment, record five lines: the decision, the claim that must be true, the observed behavior, the strongest alternative explanation, and the action the result authorizes. If the action requires a stronger signal than the observation provides, narrow the action or run another test. This small record prevents a series of individually reasonable experiments from becoming one inflated story.
Also record disconfirming behavior. Prospects who refuse to share inputs, users who reach value but do not return, and buyers who praise the outcome but decline realistic terms are not failed relationships. They locate the boundary of the evidence. The useful question is which assumption changed, not whether the experiment can still be described as positive.
Customer interest matters. It tells a founder where attention, curiosity, and willingness to engage may exist. Product evidence begins when observed behavior touches the actual value, recurrence, or transaction being claimed. Keep the conclusion no larger than the signal, and every experiment becomes more useful: not a vote on the entire company, but a disciplined reason to make the next decision.

