Not without checking. Modern models are fluent, confident, and wrong often enough that publishing their output unverified is a real risk. The danger is not that a model produces obvious nonsense. It is that it produces something plausible, well phrased, and attributed to a source that does not say what the model claims.
We use AI heavily in our own content work, so this is not a warning from the sidelines. It is the thing we have had to build an actual process around, because the failure mode is quiet. A fabricated statistic does not look like an error. It looks like the best sentence in the paragraph.
What follows is how we think about verification, what the measured error rates actually are, and where the real risk sits. Most of it is unglamorous, and that is rather the point.
A hallucination is output that is fluent and confident but not supported by anything real. The model is not lying, because it has no concept of truth to violate. It is producing the most statistically plausible continuation of your prompt, and plausible text and true text overlap most of the time but not all of it.
The important consequence is that hallucinations do not feel different from correct answers. There is no hedging, no tell, no drop in quality. A fabricated report title reads exactly like a real one because the model built it from the same patterns real report titles follow.
This is why intuition fails as a defence. People assume they will notice something wrong, and they do notice when a model claims something absurd. They do not notice when it attributes a real sounding percentage to a real company in a plausible year. That is the case that gets published.
The categories most prone to it are the ones with the most specific detail. Dates, version numbers, study names, quotations, dollar figures, and URLs. Anything precise enough to sound authoritative is precise enough to be invented.
Less often than the panic suggests, and more often than publishing unverified would justify. Vectara runs a public Hallucination Leaderboard measuring how often models introduce unsupported content when summarising a document they have been given. As of its 11 May 2026 update, the best scores sat between roughly 1.8% and 4.1%.
The methodology is worth understanding before you read anything into those numbers. Vectara uses its HHEM-2.3 detection model against a private dataset of more than 7,700 articles ranging from 50 to 24,000 words, prompting each model at temperature zero to summarise using only the information provided.
That last phrase is the whole story. This benchmark measures grounded summarisation, where the model has the source document sitting in front of it. It does not measure what happens when you ask a model to recall a statistic from memory, which is a far harder task and a far higher risk one.
So a low single digit rate on a leaderboard is not permission to trust recalled facts. It is evidence that models are good at not contradicting a document you hand them. The moment you ask for something the model has to remember rather than read, you are in different territory and the leaderboard does not apply.
No, and this is the most persistent myth in the topic. Google's position is about purpose, not production method. Its guidance on helpful content states plainly that using automation, including AI generation, to produce content for the primary purpose of manipulating search rankings is a violation of its spam policies.
Read that carefully, because the condition is doing the work. The violation is manipulating rankings, not using a tool. Google's spam policies define scaled content abuse as generating many pages for the primary purpose of manipulating search rankings and not helping users, and name using generative AI to produce many pages without adding value for users as an example.
The distinction is between AI as a writing tool and AI as a volume machine. Nobody at Google objects to a model helping you draft a genuinely useful article. What triggers the policy is producing a thousand thin pages because a model made it cheap to.
Google also gives you a concrete quality question that maps directly onto fact checking. Among its assessments of expertise, it asks whether content has any easily verified factual errors. That is a strikingly literal test, and it is the one an unverified AI draft fails. We go deeper into the trust side of this in our guide to what E-E-A-T means and how to prove it.
Anything specific enough to be quoted against you. Statistics with a source and year. Product launches, version numbers, and feature releases. Named studies and reports. Direct quotations. Company financials. And any URL, because a plausible looking URL is one of the easiest things for a model to construct.
General explanations are much lower risk. If a model explains how HTTP caching works or why line length affects readability, it is drawing on a well represented consensus, and errors there tend to be conceptual rather than fabricated. You should still know the subject, but the failure mode is different.
The highest risk category is the one that feels most credible, which is a stat attributed to a named organisation in a named year. That construction is exactly what makes writing feel researched, and it is exactly what a model can produce from thin air because it has seen ten thousand sentences shaped like it.
Our internal rule is blunt. If a sentence contains a number and a source, that sentence does not ship until someone has opened the source and seen the number. No exceptions for figures that sound obviously right, because sounding right is the problem rather than the reassurance.
Go to the primary source and read it. For a statistic, that means the organisation's own publication rather than a blog citing it. For a product feature, the vendor's own documentation or changelog. For a study, the paper or the research org's own page. A search result summarising the claim is not verification.
The distinction between primary and secondary matters more than people expect, because errors propagate. A figure gets slightly misquoted in one blog post, that post gets cited by three more, and within a year the wrong number has more search results than the right one. Models trained on that corpus will confidently reproduce the wrong version.
Retrieval is not enough on its own either. Finding a page that mentions the topic does not confirm the specific figure. You have to locate the actual number on the actual page, which is slower and is the only step that genuinely counts.
When a source cannot be found, the correct action is to cut the claim rather than soften it. Rewriting a fabricated stat into studies show is worse, not better, because it keeps the unfounded implication and removes the ability of a reader to check it. Deleting is always available and always safe.
Because it passes every check except the one that matters. A fabricated citation typically names a real company, a real report format, and a believable year. Everything about it is checkable in principle, and nothing about it looks wrong, so it survives any review that is not an actual retrieval.
This is where reputational risk turns concrete. If you publish a statistic attributed to a named company and that company never published it, their communications team can correctly say your article contains a false statement about them. That is a takedown request rather than a typo.
The same applies to product claims. Writing that a platform launched a feature it never launched is not a small error. It is a factual assertion about a company's actions, and the company is the one authority that can definitively contradict you.
The test we apply is to imagine the named organisation reading the piece the next morning and asking whether there is a single sentence they could correctly call false. If the answer is yes, the sentence changes or goes, regardless of how much it improves the paragraph.
Only after you have verified them yourself, at which point they are not AI-generated statistics any more. They are statistics you confirmed, which a model happened to point you toward. That reframing is useful, because it puts the responsibility in the right place.
Models are genuinely good at the pointing part. Asking what research exists on a topic is a reasonable use, and it will often surface real studies you would not have found. The mistake is treating the output as the finding rather than as a lead worth chasing.
The economics of this are better than people assume. Verification is fast when the source exists, because you are confirming rather than searching. It is slow only when the source does not exist, and that slowness is the signal. If you cannot find it in a few minutes, that is usually the answer.
Three verified statistics beat eight where two are invented, and not by a small margin. The eight-stat version reads as more authoritative right up until someone checks one, at which point every other number in the piece becomes suspect. Precision you cannot defend is a liability wearing the costume of rigour.
It is a hard gate rather than a review step. Nothing we publish moves forward while a claim is unverified. Every statistic must have a source we opened during that piece of work, every product or launch claim must be confirmed against the vendor's own documentation, and every external link must be a page we actually loaded.
We hold the same standard for claims about ourselves, which is the part most studios skip. It is tempting to write that we tested fourteen tools or that a client saw a specific lift, because specificity is persuasive. If we cannot support the number, it does not appear. Honest generality beats manufactured detail, and readers can tell the difference more often than writers think.
Before anything publishes, we ask three questions in writing. Does this assert any event, product, or number I did not confirm against a primary source. Could the named company correctly call any sentence false. Did I invent any detail to make the piece feel better sourced than it is. Any yes stops the piece.
Holding an article because the sources did not hold up is a success of the process rather than a failure of the day. That framing matters, because the pressure to publish is constant and the cost of one fabricated fact is far higher than the cost of a missing article. The same instinct drives how we use AI to QA a website before launch, where the tool proposes and a human confirms.
Pick one published article on your site that used AI in the drafting and check every number in it. Open each source. See whether the figure is actually there. Most people find at least one claim they cannot trace, and finding it before a reader does is the entire game.
Then write your rule down before you need it. A verification standard invented under deadline pressure is a standard that bends. Decide now that unsourced numbers get cut, and the decision stops being a judgment call every time.
Using AI to write is not the risk. Publishing without checking is. If you are building a content process and want help designing the verification step into it rather than bolting it on afterwards, we are happy to talk it through, and our take on letting AI write your website copy covers the drafting side. Come find us at phoenix.studio.
Tell us where you want to go. We'll tell you how we'd get you there.