What MiroFish cannot do
Every AI prediction product has limits, and almost none of them publish theirs. Here are ours, including the ones that are bad for us: 24 agents is a sample, not a population; the agents are language models, not people; we have never scored a prediction against what actually happened; and agreement among simulated agents is not evidence of being right.
If any of this makes the product less useful to you than you hoped, that is the correct conclusion to draw, and we would rather you drew it here than after paying.
Written by Zinedine · Founder & Developer · Published 27 August 2026
Hard limit
24 agents is a sample, not a population
A run casts 24 stakeholders from your scenario and runs them through 3 rounds. That is enough to surface the range of reactions and the arguments that spread; it is nowhere near enough to model a market, an electorate, or a customer base as a statistical population.
Treat the output as a well-briefed room arguing about your decision, not as a survey. If you need to know what 40,000 people will do, this will tell you which arguments they are likely to reach for — not how many of them will act.
Hard limit
The agents are not people, and their quotes are not evidence
Every agent is a language model given a role and an incentive. None has ever bought anything, cancelled a subscription, or been annoyed by a price rise. They are reasoning about how such a person would react, which is a different thing from being one.
The quotes in a report are generated, not collected. They are useful as a rehearsal of the objection you will face. They must never be pasted into a deck as a customer quote, cited as research, or shown to anyone who might mistake them for real.
Hard limit
We have never scored a prediction against what actually happened
This is the limitation we like least and the one you should weigh most. MiroFish reports carry probabilities, but we have no calibration data: no archive of predictions checked against outcomes, no measured hit rate, nothing that would let anyone say a "70%" here is right seven times in ten.
Until that exists, the probability figures express how the simulated room divided, not a validated forecast. Anyone quoting an accuracy rate for a tool like this — ours or a competitor’s — is quoting something nobody has measured.
Hard limit
Agreement among agents is not evidence of being right
The confidence figure on a report is real in a narrow sense: it is the share of agents that landed on the same stance in the final round, and it genuinely varies — we have measured runs from 42% to 92%. What it measures is how consistently the scenario points one way once the model is reasoning about it.
It does not measure correctness. Twenty-four agents agreeing can mean the situation is clear, or it can mean the framing you gave them made one answer obvious. A high number is a reason to ask whether you described the tension honestly, not a reason to stop thinking.
Worth knowing
It gets arithmetic wrong
Language models are unreliable calculators, and reports sometimes derive a figure incorrectly from numbers you supplied. A published example on this site computes a $6,000 loss where the correct answer is closer to $2,000, and we left it visible rather than regenerating it.
We have tightened the instructions and now verify one class of error automatically — probability sets must sum to 100%, checked against real runs before each release. Derived currency figures are not automatically verified. Check any number you intend to act on.
Hard limit
It cannot see anything you have not told it
There is no connection to your billing system, CRM, analytics, support inbox, or contracts. The simulation knows exactly what is in your prompt and the document you attached, and nothing else.
This is why vague inputs produce vague reports, and why naming two or three real specifics — an actual segment, an actual competitor, the actual number — changes the output more than any other thing you can do.
Hard limit
It cannot know what has not happened yet, or what is not public
The decisive fact in a crisis is often a document a journalist already holds and you do not. The decisive fact in a launch is often a competitor’s unannounced plan. No simulation reaches those, and a confident report that omits them is confidently incomplete.
Crisis scenarios are the weakest case for this reason, and our crisis pages say so directly: confidence figures there run lower than on pricing or launch work, and they should.
Worth knowing
Reports are written in English, whatever you write in
You can submit a scenario in any language and the engine will reason about it, but the report comes back in English. The site is available in Chinese, Spanish and French; the output is not.
Testing messaging or tone in a language other than English is also genuinely less reliable — idiom and register are where simulated interpretation is weakest. Get a native reader to confirm anything you plan to spend money on.
Worth knowing
What is actually under the hood
Runs use gpt-4o-mini. A prediction is 74 model calls: one to cast the stakeholders, 72 agent reactions across 3 rounds, and one to write the report from the transcript. It takes about 29 seconds and costs us about 1.2 cents.
We publish this because the alternative is letting you assume something grander. If a different model or a larger cast would change your decision to subscribe, you now have what you need to decide.
Worth knowing
A run can be smaller than intended, and will tell you so
If an agent call fails it is dropped; if a round runs out of time the report is written from the rounds that completed. The agent and round counts shown on a report are counted from work that actually happened, so a degraded run shows smaller numbers rather than pretending otherwise.
This is a change. Until 27 August 2026 those counts were generated randomly and bore no relationship to anything. We replaced the engine rather than the labels.
Hard limit
It is not professional advice
Nothing a report says is legal, financial, medical, or employment advice. Where liability, regulatory disclosure, securities, or employment law is in play, a simulation is at most an input to a conversation with someone qualified and insured.
It is also not a substitute for talking to real customers. If you can ask ten people who actually pay you, do that first — it beats any simulation, including this one.
What it is genuinely good for
Having read all of that, the honest positive case: it is a fast, cheap way to be argued with by something that has no stake in your decision and no reason to be polite about it. It surfaces the objection you had not thought of, the segment that reacts in the opposite direction to the one you were worried about, and the second-day story you were not planning for — while changing course is still cheap.
That is worth about thirty seconds and a few cents. It is not worth mistaking for a forecast.
Questions
Are AI predictions accurate?
Nobody selling one can honestly tell you, including us. MiroFish has no calibration data — we have never scored an archive of predictions against real outcomes, so the probability figures on our reports express how the simulated stakeholders divided rather than a validated hit rate. Any vendor quoting an accuracy percentage for this class of tool is quoting a number that has not been measured.
How many agents does MiroFish actually run?
24 per run, across 3 rounds of reaction — 72 agent responses in total, plus one call to cast the stakeholders and one to write the report. That is a sample of perspectives large enough to surface the range of reactions and the arguments that spread, and far too small to model a market as a population.
Does a high confidence score mean the prediction is right?
No. The confidence figure is the share of agents that reached the same stance in the final round — real, measured, and varying between runs from about 42% to 92%. It measures how consistently the scenario points one way, not whether the answer is correct. A high figure can equally mean your framing made one answer obvious.
Can I use the quotes from a report as customer research?
No. They are produced by simulated agents, not collected from people. They are useful for rehearsing objections you will face and must never be presented as customer quotes, cited as research, or shown anywhere they could be mistaken for real feedback.
Can MiroFish see my company data?
No. It has no connection to your billing system, CRM, analytics, or contracts. It knows only what is in your prompt and any document you attach. That is why naming a few real specifics improves the output more than anything else you can do.
What is MiroFish worst at?
Crisis prediction, because the decisive variable is usually a fact that is not yet public — often something a journalist already holds. Long-horizon reputational change is second, since effects unfolding over years are beyond what a few rounds of simulation can meaningfully estimate. Both are stated on the relevant pages rather than only here.
Found a limitation we have not listed, or a claim elsewhere on this site that contradicts something here? Tell us and we will fix the page rather than the complaint. See also a real report with its errors annotated and our published research.
Still useful to you? See exactly what it produces before paying anything.
Read the example report