Search for "AI assistants" and you'll get chatbots, copilots, agents, and "digital employees" all sold as the same thing. Then you pick one, ask it to book a meeting, and it politely explains how meetings work. Or worse: you pick an "agent," and it emails a client before you've read the draft.
My name is Artem, and I run the Writingmate blog. I've spent the last couple of years switching between models all day, and the single most useful thing I learned wasn't which model is smartest. It was how much freedom to give the thing. This post gives you a simple way to decide that, ten real tasks mapped to it, and a test you can finish in ten minutes.
Short version: stop asking "assistant or agent?" and start asking "how much should this do without me?" The answer changes by task, not by brand.
The vocabulary, without the marketing
Every vendor defines these words slightly differently, and the definitions overlap. Even people who write about this for a living disagree. But the dividing line most explainers land on is the same: who decides the next step, you or the software.
- Chatbot: talks. You ask, it answers, usually from its training or a fixed knowledge base. It doesn't touch anything outside the chat window.
- AI assistant: helps you do work. It drafts, summarizes, searches, and may read your files or calendar, but you still decide and press the button.
- AI agent: takes a goal, picks the steps, uses tools, and acts. Anthropic's engineering team draws the line cleanly: in a workflow, your code fixes the path; in an agent, the model directs its own process and tool use (see their Building effective agents post). They also advise using the simplest option that works, which sometimes means no agent at all.
And the second keyword people trip on, generative vs agentic AI? Generative AI produces content: text, images, code. Agentic AI uses that same generation ability inside a loop of plan, act, check, repeat. Same engine, different amount of steering wheel.
The autonomy ladder: four rungs
This is the model I use instead of the three labels. Each rung gives the AI one more permission.
Rung | What it does | Who acts | Typical failure |
|---|---|---|---|
1. Answers | Explains, looks things up, compares | You | Confident wrong answer |
2. Drafts | Writes emails, docs, code, plans for you to review | You | Generic or off-tone output |
3. Acts with approval | Prepares an action (send, book, update) and waits for a yes | AI proposes, you confirm | Rubber-stamping approvals |
4. Acts unattended | Runs on a trigger or schedule, no one watching | AI | Silent errors and runaway cost |
Chatbots live on rungs 1 and 2. Most assistants cover 1 to 3. "Agent" usually means rung 3 or 4. Notice the failure column: the higher you climb, the less visible mistakes become. That's the real cost of autonomy, and it's why you shouldn't climb higher than a task earns.
Ten common tasks, mapped to a rung
Here's where I'd put ten tasks I see people hand to AI all the time, with a starting prompt for each. Adjust the details; the structure is what matters.
Task | Rung | Example prompt |
|---|---|---|
Explain a contract clause | 1 Answers | "Explain this clause in plain English and list two questions to ask a lawyer." |
Compare two tools | 1 Answers | "Compare X and Y for a 5-person team. Table: price, limits, gaps. Cite sources." |
Write a client email | 2 Drafts | "Draft a polite reply declining the scope change. Under 120 words, firm but warm." |
Summarize a long report | 2 Drafts | "Summarize in 8 bullets, then list what the author didn't prove." |
Debug a function | 2 Drafts | "Find the bug, show the minimal fix, explain why it failed." |
Plan a trip or week | 2 Drafts | "Build a 3-day itinerary with travel times. Flag anything that needs booking." |
Reply to routine inbox mail | 3 With approval | "Draft replies to these 10 emails and label each: send as-is, edit, or needs me." |
Schedule meetings | 3 With approval | "Propose three slots from my calendar and write the invite. Don't send." |
Update CRM records after calls | 3 With approval | "From these notes, list the field changes you'd make. Wait for my OK." |
Weekly competitor digest | 4 Unattended | "Every Monday, check these five pages, summarize changes, post to the team channel." |
Only one of the ten earns rung 4, and that one is read-only: it gathers and reports, it doesn't change anything. That's a pattern worth stealing. Unattended is safest when the worst outcome is a mediocre summary.
How I tested: the same task, three models, ten minutes
Here's the method I use before I trust any assistant with a repeat job. It takes about ten minutes and works with any tool where you can switch models. In Writingmate, I open the chat and run the same prompt against three different models, because a model that's great at drafting is often mediocre at following a strict checklist.
- Pick one real task you did this week, not a toy example. Paste in the real inputs (strip anything private).
- Write the prompt once with a clear output format: "table with three columns," "under 150 words," "label each item."
- Run it on three models: one fast and cheap, one flagship, one from a different vendor. Don't change a word between runs.
- Score each output 1 to 5 on accuracy, format-following, tone, and "would I send this untouched?"
- Re-run the winner once. If the result changes a lot between two identical runs, that task isn't ready for rung 3 or 4.
The last step is the one people skip, and it's the most revealing. Consistency is what you're really buying when you automate. In my runs, the cheaper model often ties the flagship on routine formatting jobs and loses on anything that needs judgment, which is why I wouldn't pick one model for everything.
This is the practical argument for a multi-model setup over a single-brand assistant: you can run that comparison in one tab instead of juggling subscriptions. If you're pricing that out, see Writingmate pricing.
What people say when agents go wrong
The honest case against jumping to rung 4 comes from people who did it. A thread on r/AI_Agents asked what the worst quiet cost of an agent had been, and one reply sums up the failure mode:
"[Agents] fail quietly, by spending your money while you sleep... a slow drip that had turned into £220 by the time I caught it." — r/AI_Agents on Reddit
And on the other side, here's a user comparing assistant behavior with agent behavior on the same request:
"Ran the same prompt through Grok and GPT side by side for images. Grok nailed it in one shot, GPT needed three tries. But GPT's agent mode actually finished my task end to end and Grok's just... talked about it." — @aivisionary on X
Both are true. Agents finish things that assistants only talk about, and agents can also finish things you didn't want finished. The ladder exists so you choose which of those you're signing up for.
The approval checklist before an AI takes any action
Before you move a task from rung 2 to rung 3, or rung 3 to rung 4, I'd run through these questions. If you answer "no" to any of the first four, stay one rung lower.
- Is the action reversible? Drafting a calendar hold is; sending an email to a customer isn't.
- Can I see exactly what it will do before it does it? If you can't preview, you can't approve.
- Does it have the minimum access? Read-only first. Add write access to one folder, one inbox label, one project, not everything.
- Did it pass the repeat test? Three identical runs, three acceptable results.
- Is there a spending or step limit? A cap on runs, tokens, or actions per day stops a quiet drip becoming a bill.
- Is there a log? You should be able to answer "what did it do last Tuesday?"
- Who gets notified when it fails? Silent failure is the one that costs you.
For anything touching money, legal commitments, customer data, or public posting, I keep a human approval step permanently. Not because the models are bad, but because the cost of one wrong action dwarfs the minutes saved.
From chat to a custom agent: a short path
Here's how a repeat task graduates. I'll use "weekly competitor digest" as the example, but any recurring job works.
- Do it in chat three times (rung 1 to 2). Save the version of the prompt that worked best.
- Turn that prompt into standing instructions: role, inputs, output format, what to do when information is missing, what never to do.
- Pick the model using the ten-minute test above, not the name on the box.
- Add only the tools it needs, such as web search or a file source, read-only at first.
- Run it with approval on for two weeks (rung 3). Track how often you edit its output.
- Remove approval only when edits drop to near zero, and keep it for anything irreversible.
Writingmate lets you do this without code. The walkthrough is in the docs on creating custom agents, and if you want a longer worked example, my earlier post on building an AI agent from goal to tested workflow covers instructions, tools, and a five-case test.
So which one do you need?
Use this as a quick filter:
- You mostly ask questions or want writing help: a chatbot or assistant (rungs 1 to 2) is enough. Paying for agent features is waste.
- You do the same multi-step task every week: build a rung 3 assistant with approval. This is where most of the time savings live.
- The task is read-only, repeated, and low-stakes: rung 4 is fine.
- The task spends money, contacts customers, or deletes things: keep a human in the loop, whatever the vendor calls it.
The label on the product matters far less than the permissions behind it. A "chatbot" with calendar access is an assistant, and an "agent" with no tools is a chatbot with better branding. Check what it can touch, test it on a real task, and climb the ladder one rung at a time.
See you in the next one!
Artem
Frequently Asked Questions
Sources
Written by
Artem Vysotsky
Ex-Staff Engineer at Meta. Building the technical foundation to make AI accessible to everyone.
Reviewed by
Sergey Vysotsky
Ex-Chief Editor / PM at Mosaic. Passionate about making AI accessible and affordable for everyone.
