The ability of AI agents to carry out concrete actions has produced expectations about their autonomy that are sometimes excessive. In practice there are real limits worth knowing, if you want to design systems that are dependable rather than systems that promise more than they deliver.
Very long multi-step tasks
The longer a sequence of steps to be carried out without human involvement, the greater the chance that an error in an intermediate step propagates through the rest unnoticed. For anything critical, breaking the sequence with checkpoints remains more reliable than entrusting everything to one long autonomous run.
Ambiguity in the original request
An AI agent interprets natural language, but ambiguous requests can lead to plausible yet wrong interpretations. A good agent recognises the ambiguity and asks rather than proceeding on an unverified assumption — but that behaviour has to be designed in deliberately; it is not guaranteed by default.
Oversight is still needed
For actions with financial or irreversible consequences, human oversight — even just an explicit confirmation before proceeding — remains the most reliable way to use an AI agent today. This is not a temporary limitation to be removed as soon as possible, but a prudent design choice for any system touching real company data.