How do I prompt an AI to say 'I don't know'?
Give clear permission
Models often guess because they are trained to be helpful and to produce an answer. Saying 'It is fine to say I don't know' removes that pressure.
Be specific about when to use it. A rule like 'If the source text does not contain the answer, reply exactly: I don't know' is much clearer than 'Be honest if you are unsure.'
You can also ask for a short confidence note, such as 'low, medium, or high,' next to each claim. That gives you a warning sign without forcing a yes-or-no answer.
Make it easy to comply
Supply the source material in the prompt and tell the model to use nothing else. When the model has a defined pool of information, 'I don't know' becomes a real option instead of a failure.
Set a fallback for partial answers: 'If you know part of it, answer that part and mark the rest as unknown.' This avoids all-or-nothing replies.
Test the rule with a question you know is not in the source. If the model still invents an answer, tighten the wording and repeat the rule at the end of the prompt.
- Say: 'It is acceptable to say I don't know.'
- Add: 'Answer only from the text below.'
- Add: 'If the answer is not present, reply I don't know.'
- Ask for a confidence level with each claim.
- Test with a question the source cannot answer.
Common mistakes
- Assuming a polite instruction like 'be accurate' is enough; you need a specific trigger and fallback wording.
- Punishing uncertainty in follow-ups, which teaches the model to guess instead.
- Forgetting that 'I don't know' still needs checking; the model may say it when the answer was actually available.
