I asked ChatGPT for a list of the latest academic papers on AI and consumer psychology last week.
It gave me five citations, complete with authors, journals, and publication dates.
The titles sounded perfect.
The authors are real people who study related topics.
The journal names were legit.
Everything looked right, until I tried to verify them.
One of them doesn’t exist.
It was completely made up, yet presented with total confidence.
This is what AI hallucinations look like in practice.
Why AI Makes Things Up
It helps to understand what causes hallucinations, because it changes the way you use the tools.
At the most basic level, language models generate responses by predicting likely words and phrases based on patterns they learned from massive amounts of training data, like advanced versions of autocomplete or highly sophisticated “language prediction machines.”
That’s why they can sound extremely convincing even when they’re wrong.
Newer reasoning models can “think” through problems more deliberately, which is why they hallucinate a lot less.
So why do they still happen?
One of the clearest explanations I’ve seen came from OpenAI researchers last year.
They argued that current training and evaluation methods often optimize models to be good test-takers, rewarding them for guessing over acknowledging uncertainty when they’re not sure about the answer.
This encourages models to produce plausible-sounding answers instead of saying “I don’t know” to improve their test performance.
But there are other reasons, too.
Hallucinations are more likely when:
the model is missing key information
the sources it’s pulling from conflict with each other
the prompt is vague
facts have changed since training
The good news is that this is getting better as the tech advances.
Leading AI models today hallucinate far less than they did even a year ago, especially as they’ve become better at reasoning, searching the internet, and working more accurately with grounded material.
Better, though, is not the same as reliable.
If you’re using AI for high stakes situations that require accuracy or precision—strategic or financial analysis, research, legal work, medical advice—even minor inaccuracies make the output unreliable and potentially harmful.
So, what can you do?
How to Reduce AI Hallucinations
You can’t get rid of hallucinations completely.
You can make them much less likely.
The key is giving the model less room to guess.
That usually means being clearer about three things: what it should use, what it needs to verify, and what it should do when it isn’t sure.
You can do that by adding specific instructions to your prompts.
The examples below are simple on purpose. Think of them as starting points.
I usually combine several of these in the same prompt and customize the instructions to the specific task, project and stakes.
👉 One important note before we get into it: if you’ve been reading me for a while, you already know what I’m about to say, but it bears repeating.
Use a paid plan.
The most advanced reasoning models—especially with extended thinking turned on— are smarter and hallucinate far less.
If you’re using AI for anything beyond casual chat, the $20/month tier on ChatGPT or Claude should be the bare minimum.
Ground Responses in Your Sources
For tasks when accuracy matters, do not start with a blank canvas.
Give the model the documents, data, transcripts and notes to use and explicitly tell it to only use what you’ve provided.
👉 Prompt addition:
“Use only the sources I’ve provided. If the materials don’t support clear answers, say so rather than guessing or filling in the gaps.
If a question requires information not found here, list the specific ‘missing data points’ I need to provide for you to give me reliable answers.”
Give It Permission to Say “I Don’t Know”
This one might seem too simple, but it works. Welcoming uncertainty stops AI from guessing when it doesn’t have the required context.
👉 Prompt addition:
“If you’re not confident in any part of your answer, flag it. I’d rather get a partial answer with clear caveats and know you’re uncertain about other parts than get a confident-sounding guess.”
Or
“For each claim you make, rate your confidence: high, medium, or low, and explain what you’re uncertain about.”
Ask for Sources—Then Verify Them
Requesting citations forces the model to anchor claims to real information. But AI can still invent sources.
That means asking for sources is only step one.
You still need to check them.
👉 Prompt addition:
“Provide a source for every factual claim. Include a direct link. If a claim can’t be verified, label it unverified instead of guessing.”
Or
“Before answering, verify claims against credible sources. Use primary sources when possible, and for each factual claim, show the source it came from. If you cannot verify something, flag it clearly instead of guessing.”
And then:
“Before you finalize, audit every citation and URL. Confirm that the source exists, the title, author, date, and publication details are correct, the link goes to the right page, and the source actually supports the claim attached to it. If any part is unclear, label it unverified.”
Yes, you’re asking the model to check its own work. It’s not foolproof, but it catches more stuff than you’d expect.
Ask It to Use Web Search
The latest set of reasoning models will often use search automatically when your query seems to need current information. But don’t leave that call up to them.
If you’re asking AI about anything that requires up-to-date factual information, explicitly ask it to use web search, pull from primary or otherwise credible sources, and cite what it used.
👉 Prompt addition:
“Use web search for this. Pull from primary or otherwise credible sources, and cite the sources you used for each key claim. If anything can’t be verified, flag it instead of guessing.”
When Hallucinations Are Useful
As I’ve written before, AI hallucinations are a double-edged sword.
While they pose challenges in areas that demand accuracy and precision, they’re genuinely valuable when creativity is the goal.
The very thing that makes AI unreliable for fact-based tasks makes it surprisingly powerful for brainstorming, ideation, and creative problem solving.
This unpredictability allows AI to make complex and unexpected connections that humans might overlook.
Think of it like a brainstorming session—most ideas go nowhere, but one or two wild ideas can spark breakthroughs.
That’s why scientists are embracing hallucinations as a creative tool.
They’re using them to create new proteins that don’t exist in nature, speed up drug discovery, and design innovative medical devices.
So far, these hallucinations have led to new cancer treatments, therapies for viral infections and a new catheter with unique grooves to reduce bacterial infections—a simple but powerful fix for a widespread problem.
If you want to push AI toward more original and creative responses, check out this edition where I share a prompting technique that does exactly that.
So, what does this mean for us?
AI hallucinations aren’t going away. Not completely.
Researchers continue to work toward models that are more self-aware, teaching them to recognize when they don’t actually know the answer, admit uncertainty, and even flag when an answer might be unreliable.
Our job is to continue to build and deepen our instinct on how to best use these models based on our goals.
🛡️ When accuracy and precision are crucial, use the guidance above to write specific and detailed instructions and review and verify the outputs.
💡 When creativity and exploration are the priority, embrace hallucinations to fuel more fresh, distinct and unconventional ideas, and then refine the most promising suggestions with our own expertise and judgement.
This will help you better manage the risks while taking full advantage of the times hallucinations might actually lead you somewhere interesting.
P.S. Share this with your team.
Most corporate AI trainings cover hallucinations in a few slides—usually some version of “verify everything”—without explaining why they happen or practical and actionable ways to reduce them.
Strong prompting is still the foundation, but the additional methods I shared here add a layer even experienced power users haven’t learned yet.
That’s a real risk. And it grows every day as more of your team uses AI for more of their work.
The new reasoning models make so much more possible, but only if you know how to use them.
📈 Work with Me
AI Advising and Consulting: Strategic guidance to build sustainable AI strategies and adapt to shifting audience behavior
AI Training: Practical training that helps teams use AI more strategically, effectively, and safely—so it becomes a trusted thought partner, not just a task assistant
What You Need to Know About AI This Week ⚡
Clickable links appear underlined in emails and in orange in the Substack app.
🎭 AI companies are now recruiting improv actors to train their models.
Handshake AI, a company that supplies training data to OpenAI and other major labs, is hiring improv actors, sketch comedians, and theater performers to help train AI on human emotion and interaction.
The listing calls for people with “emotional awareness,” the “ability to recognize, express, and shift between emotions in a way that feels authentic and human” and “interactions that feel grounded, human, and fun to play.”
Participants would improvise unscripted, open-ended scenes over video and get paid around $74/hour.
Training data companies like Handshake have already been hiring professionals—lawyers, doctors, accountants, management consultants—to help AI get smarter in those fields.
This time, what they seem to want is not subject-matter expertise, but better examples of human behavior and interaction.
That makes sense.
AI models are often described as “jagged,” meaning they can be surprisingly good at complex tasks while failing at simple human ones.
The latest reasoning models can handle advanced research and analysis, but still struggle with social dynamics and sound robotic in ways that are hard to pin down—which is what made those Anthropic Super Bowl commercials so funny.
Here’s a funny example of that broader problem. Someone asked ChatGPT 5.2 whether to walk or drive 100 meters to the car wash.
This was the response 🤣:
Handshake hasn’t said what this data will be used for.
But the focus on emotional range and spontaneity suggests the goal is to make AI interactions feel more natural and more human.
I’m still thinking through what this new training data will make possible, but the bigger pattern here is that demand for niche and specialized training data continues to surge.
🛑 Hollywood has put China’s hottest AI video model on pause.
ByteDance has paused the global rollout of Seedance 2.0, the model behind those fake Tom Cruise–Brad Pitt clips that went viral after its China launch in February.
The videos triggered backlash from Hollywood studios, which sent cease-and-desist letters and accused ByteDance of using copyrighted material without permission.
The company had been targeting a worldwide launch in mid-March but is now delaying it while its engineers and legal teams work on stronger copyright protections.
✍️ Should we care if humans prefer AI writing?
The New York Times recently published a blind quiz asking readers to choose between human and AI-written passages.
More than 86,000 people took it, and the results were narrower and more revealing than most people expected: readers slightly preferred the AI writing overall, 54% to 46%.
I took the quiz too, and in my case the humans won four out of five times.
Reid Hoffman wrote a thoughtful response to the whole thing, and his argument is basically this: the results say less about AI replacing human expression and more about what people often want from writing in the first place.
A huge amount of writing is not art. It is practical.
It’s there to explain, summarize, clarify, guide, or answer questions.
In those contexts, we’re usually not looking for originality or soul.
We want something clear and useful.
AI is often very good at exactly that kind of writing.
“Let’s pause to ask: is this necessarily a bad thing?
In many cases, it is also an opportunity to move human effort up the value chain away from routine drafting and toward originality and taste.
The writing work that remains most valuable may become more distinctly human, not less.”
So no, this does not mean human expression is obsolete.
It means AI may take over more of the routine, functional writing and leave more room “for the kind of writing only humans can really own: the kind that carries judgment, originality, courage, and soul.”
That feels pretty right to me…
🤖 Why OpenClaw is taking off in China.
China isn’t just embracing OpenClaw, the popular open-source AI agent that can take actions across apps and devices.
It’s trying to build an economy around it.
I wrote about OpenClaw a few weeks ago, when it was still mostly a buzzy AI story.
Since then, OpenAI hired its founder, and in China OpenClaw has grown into something much bigger: a consumer phenomenon AND a massive business opportunity.
That’s because it serves several agendas at once.
1️⃣ Chinese cloud companies think—Amazon Web Services or Google Cloud, but in China—want the extra AI demand, since agents use far more processing power than ordinary chatbot use.
2️⃣ Tencent, the Chinese tech giant behind WeChat, sees agents as a chance to compete through its existing apps, distribution, and user experience, not just model quality.
3️⃣ Local governments see OpenClaw as a possible engine for new startups and “one-person companies,” where AI can do the work that previously required entire teams.
But agents with broad permissions can leak, misuse, or delete sensitive data, which is why Chinese regulators are promoting the upside while also warning about the major risks.
Basically, China is trying to make agentic AI useful at scale before it has fully figured out how to govern it 😨.
Meanwhile…
--
📲 Claude Cowork just got one of OpenClaw’s most useful features.
Cowork’s new Dispatch feature lets you keep one single continuous session running on your computer, message it from your phone, and come back to finished work.
That is a big part of what made OpenClaw feel so useful in the first place: the agent keeps working in the background while you check in and keep steering it from wherever you are.
I’ve been covering Cowork extensively here over the past few months, because it’s one of the clearest examples of how work is transforming, especially for non-techie professionals.
Cowork runs directly on your desktop with access to a folder on your computer. It reads your files, works across documents, and executes multi-step tasks.
It’s great for pulling information from multiple sources, making sense of it, and producing a complex deliverable.
It can also multitask, tackling different parts of a project at the same time.
If you missed my deep dive on Claude Cowork, catch up here (make sure to scroll down to the bottom of the weekly updates section.)
And if you lead a team or influence AI decisions, you need hands-on experience with tools like this NOW to understand how they could reshape your team/company’s work, especially as OpenAI expands its Codex app with similar capabilities.
🎟️ OpenAI is turning ChatGPT into a festival concierge at Lollapalooza Brazil.
As part of its first big marketing push in Brazil, OpenAI is sponsoring the festival and launching a custom GPT to help attendees plan their day, answer questions about programming, and navigate the event with prompts like “organize my perfect day at Lolla 26.”
This is a really smart example of how AI can be used in practical and meaningful ways to enhance the audience experience.
👓 How private are Meta’s AI glasses, really?
Meta is facing a class action lawsuit after a joint investigation found that sensitive footage from Ray-Ban Meta glasses—including nudity, sex, bathroom visits, and bank card details—was routed to contractors in Kenya who review and label clips to help train Meta’s AI.
The suit says Meta’s privacy claims misled buyers about its actual review process for the footage.
This is the uneasy truth about AI wearables: they’re pushing surveillance and consent into riskier territory long before most people understand the trade-offs.
🍸 Bravo fans are getting an AI Andy Cohen.
Peacock is launching a personalized Bravo feed in its mobile app, built around your favorite shows and personalities, with an AI-powered Andy Cohen guiding the experience, recommending what to watch next, and even popping up to answer questions about what’s happening on screen.
This could be a very clever and fun use of AI (the personalization part is a no-brainer), but it will all come down to execution, and whether it genuinely upgrades the fan experience rather than feeling like just one more gimmick.
📊 If Excel is part of your job (my condolences), make sure you’re up to date on what ChatGPT and Claude can do there now.
Their new Excel capabilities are a big step up from where AI in spreadsheets was even a few months ago.
ChatGPT for Excel: https://chatgpt.com/apps/spreadsheets/
Claude for Excel: https://claude.com/claude-for-excel
In case you missed the last edition, you can find it 👇:
🤓 ChatGPT and Claude Just Got More Useful for Real Work
Oh, it’s really hard to know where to start this week.
That's all for this week. See you in 2 weeks.
Thoughts, feedback and questions are always welcome and much appreciated. Shoot me a note at avi@joinsavvyavi.com.
Stay curious,
Avi
💙💙💙 P.S. A huge thank you to my paid subscribers and those of you who share this newsletter with curious friends and coworkers. It takes me about 20+ hours to research, curate, simplify the complex, and write this newsletter. So, your support means the world to me, as it helps me make this process sustainable (almost 😄).









I am very grateful for your support. Mark Kabakov
Yeah the whole lobster training craze going on in China is interesting. You have all of these kids and retirees allowing these agents all kinds of access. From everything that I've been reading there are a lot of security risks.