Here's the question nobody answers on a "best AI agents" page: when you walk away from your desk, does the agent come back with a finished job, or with a half-done mess and a confident summary? Most lists rank by brand and demo polish. I'd rather rank by the job, because an agent that's great at fixing a bug can be useless at triaging an inbox.
My name is Artem, and I run the Writingmate blog. I spend a lot of my week comparing AI tools and building small agents on top of different models. One caveat up front: I'm not going to hand you a fake leaderboard with invented percentages. What I'll give you is a job-by-job ranking built from public claims I could verify, plus a five-task scorecard you can run on any agent in an afternoon, so the numbers you act on are your own.
If you want the background first, this talk from Andrej Karpathy on why agents are a decade-long project is the most sober take I've found. It's a good dose of reality before you buy anything.
Agentic AI vs. generative AI, in one minute
Generative AI answers a prompt. You ask, it writes, you read. An agent takes a goal, picks tools (search, a browser, a code runner, your inbox), takes several steps, checks its own work, and returns a result. Generative AI is the engine; the agent is the engine plus a steering wheel and a task list.
That difference matters for buying decisions. A chatbot fails visibly: you read a bad answer. An agent fails quietly: step four goes wrong, steps five through nine build on it, and the final report looks fine. So the questions to ask are different. Not "how smart is it?" but "how far can it run before it needs me, and what happens when it's wrong?"
"Feels a bit like the wild west of early computing, with computer viruses (now = malicious prompts hiding in web data/tools), and not well developed defenses" — @karpathy on X
That's Karpathy on prompt injection, and it's the best argument for oversight. The moment an agent reads untrusted web pages or emails and can also send or change things, you have a risk you need to plan for, not just a quality question.
How I score agents: the five-task test
Every agent I evaluate gets the same five jobs. Use them yourself and you'll quickly see which agents deserve your trust.
- Competitor research brief. "Find five competitors to [your product], summarize pricing and positioning, cite every claim."
- Bug fix from a ticket. Give it a real, small ticket and a repo. Success means a passing test and a reviewable diff.
- Multi-step web lookup. Something that needs several pages: "Find the cheapest plan that includes SSO across these four vendors."
- Inbox triage. Sort 50 messages into reply, delegate, ignore, and draft the replies for the first group.
- Weekly report. Pull numbers from a source, compare to last week, write a one-page summary.
For each run, record four things:
- Completion without intervention: did it finish, or did you have to nudge it? Count the nudges.
- Cost per run: credits, tokens, or minutes of compute. Write the real number down.
- Failure mode: stalled, looped, hallucinated a source, or "succeeded" wrongly. The last one is the dangerous one.
- Oversight needed: none, spot-check, or review every output.
Run each task three times, not once. Agents are inconsistent, and one lucky run tells you nothing. If an agent finishes three out of three on task two, that's a signal. One out of three means it's a drafting assistant, not a worker.
Best AI agents by job
The picks below come from published roundups, such as Rework's ranking of autonomous agents by how far they run without you, and from how each vendor describes its product. I haven't re-run each one on my five tasks, so treat these as starting candidates for your own scorecard, not verdicts.
Job | Agents commonly recommended | What the claim is | What to verify |
|---|---|---|---|
Coding from tickets | Devin; Claude Code | Reads a ticket, plans, writes code, runs tests, opens a PR | Does the test pass on the first PR? Is the diff small enough to review? |
General research and tasks | Manus; Genspark | Browses, writes and runs files in a sandbox for multi-step jobs | Are citations real? Does cost spike on long runs? |
Browser workflows | Skyvern | Handles web flows where no API exists | What happens when a page layout changes? |
Recurring ops and triage | Lindy | Runs on a trigger rather than a single prompt | Does it act or only draft? Who approves sends? |
Anything repeatable you can define | A custom agent with fixed instructions | Does one narrow job the same way every time | Does your instruction set cover the edge cases? |
A pattern shows up quickly: coding agents are the most autonomous because code has a built-in judge. Tests pass or they don't. Research and ops agents have no such judge, which is why they need more oversight. A report can read beautifully and be wrong.
A warning from the community side, too. People in communities like r/AI_Agents regularly point out that many products sold as "agents" are really fixed automations with a chat box on top. That isn't bad. A fixed workflow is often more reliable than a free-roaming agent. It just means you shouldn't pay agent prices for it.
Which jobs finish unattended, and which need a human
This is my working rule after building a few of these, and it matches what the scorecard tends to show. Sort jobs by how checkable the output is:
- Checkable by a machine (run unattended): bug fixes with tests, data pulls that must match a total, format conversions. Let these run, review the result.
- Checkable by a skim (spot-check): competitor briefs with cited links, weekly reports with numbers you can compare to the source. Open three citations, confirm the numbers, done.
- Checkable only by judgment (human in the loop): inbox replies, anything customer-facing, anything that sends, spends, or deletes. Draft automatically, approve manually.
That third category is where most "fully autonomous" promises fall apart. If sending the wrong email costs you a client, an agent that drafts and waits for your click is the better product, even if it looks less impressive in a demo.
Ready-made agent vs. custom agent
The honest answer is that you'll probably use both. Here's how they compare:
Factor | Ready-made agent | Custom agent |
|---|---|---|
Setup time | Minutes | An hour or so to write instructions and test |
Best for | Open-ended jobs: "go figure this out" | Repeat jobs with a known shape |
Cost predictability | Often usage-based, can spike on long runs | You choose the model, so you control the spend |
Consistency | Varies run to run | Higher, because instructions and format are fixed |
Data control | Depends on vendor | You decide which files and integrations it sees |
The "when to build your own" rule
Build your own when all three are true: you do the job at least weekly, you can write down what a good result looks like in under ten lines, and the ready-made tool either costs more per run than you'd like or gets the format wrong. If it's a one-off, use a ready-made agent. If the job is open-ended and unpredictable, use a ready-made agent. If it's the same Monday-morning report every week, build it once.
How to build an AI agent without code
If you've been searching for how to create or make an AI agent, here's the short, no-framework version. Pick the winning workflow from your scorecard, usually the weekly report or the competitor brief, and build it in Writingmate's custom agents. According to the docs, the steps are:
- Open Agents in the sidebar and choose Add Agent.
- Fill in the name and a one-line description of the purpose.
- Write the instructions: the role, the exact output format, what to do when data is missing, and what it must never do (for example, "never invent a source").
- Pick the model. Availability depends on your plan and workspace.
- Set temperature: the docs describe 0 as consistency and 1 as creativity. For reports and triage, stay near 0.
- Add knowledge files (up to 10 MB each) such as your product positioning or last week's report, and enable the integrations or MCP tools it needs.
- Create the agent, test it with representative requests, then edit from the Agents page.
Picking the right model per step
Here's a detail worth being straight about: per the docs, each custom agent uses one selected model. So "the right model per step" means splitting the workflow into small agents, one per step, each with a model that fits. For a weekly competitor report, that might look like:
- Step 1, gather: an agent with a fast, cheap model that pulls and lists facts with links.
- Step 2, analyze: an agent with a stronger reasoning model that compares the facts and flags what changed.
- Step 3, write: an agent with a model you like for prose, instructed to use only the facts from step 2.
You paste each output into the next, or keep one chat per step. It's a little manual, but it keeps the expensive model on the one step that needs it, and it makes failures easy to find: you can see exactly which step went wrong. You can browse the available models to match each step, and our earlier walkthrough on building a weekly competitor-research agent covers the instructions in more detail. Check pricing for what your plan includes.
Oversight checklist before you let an agent run alone
- Read-only first. Let it read and draft for a week before it can send, spend, or delete.
- Require sources. Put "cite a URL for every claim; write 'not found' if you can't" in the instructions. It cuts hallucinated facts noticeably.
- Isolate untrusted content. An agent reading web pages or inbound email shouldn't also hold access to private files and outbound sending. That combination is the risk Karpathy pointed at.
- Spot-check on a schedule. Open three outputs a week. When quality slips, you'll notice early.
- Keep a rollback. Anything that edits something should leave the original recoverable.
What I'd do this week
Pick one of the five tasks, the one you actually repeat, and run it through two ready-made agents and one custom agent, three runs each. Write down completion, cost, failure mode, and how much you had to babysit. After one afternoon you'll know more than any ranking can tell you, including this one.
My recommendation: use ready-made agents for open-ended, one-off jobs and for coding with tests. Build a custom agent for anything weekly and predictable, and keep a human click on anything that leaves your outbox. If you want to try the build route, the custom agents guide takes you through it step by step.
See you in the next one!
Artem
Frequently Asked Questions
Sources
Written by
Artem Vysotsky
Ex-Staff Engineer at Meta. Building the technical foundation to make AI accessible to everyone.
Reviewed by
Sergey Vysotsky
Ex-Chief Editor / PM at Mosaic. Passionate about making AI accessible and affordable for everyone.

