Every AI agent platform demo looks the same: a friendly chat box, a few connected apps, and a task that magically finishes itself. Then you try it on your own work and the agent invents a competitor, ignores half your instructions, or burns through credits on a loop. The demo wasn't lying. It just didn't show you what to check before you commit.
My name is Artem, and I run the Writingmate blog. I spend a lot of my week building small agents for our own team and comparing how different models behave inside them. This guide is the short version of what I look at before trusting any platform, followed by a full build of one real agent: a weekly competitor-research agent you can set up without writing code.
Here's the plan. First, a five-point checklist for judging an AI agent platform. Second, a walkthrough of building the agent in Writingmate. Third, a plain-English answer to agentic AI vs generative AI, and the failure modes that sink most first agents.
How I compared platforms (and what I didn't claim)
Quick honesty note on method. I judged platforms against the same five questions, using official documentation and my own setup work in Writingmate. I'm not publishing benchmark numbers or "agent A beat agent B by 12%" figures here, because I haven't run a controlled test that would support them. The video above is a general no-code build walkthrough I'm embedding as a visual reference for the overall flow. I didn't pull specific numbers or claims from it.
What I can tell you is what repeatedly goes wrong when people choose a platform from a feature list, and the checklist below is built around those failures.
The five-point checklist for choosing an AI agent platform
Skip the logo wall and ask these five things. If a vendor can't answer one clearly, that's your answer.
- Model choice per step. Can you pick the model, and can you change it later? A research step, a summarizing step, and a drafting step don't need the same model. Locked-in platforms make you pay flagship prices for everything.
- Tool and knowledge access. What can the agent actually read and call? Look for file upload, web access, and a standard way to connect tools such as MCP servers. More important: can you turn tools off?
- Approvals. Does the agent act on its own, or does it ask before sending, deleting, or publishing? For a first agent, you want drafts and suggestions, not autonomous actions.
- Cost control. Can you see what a run costs, cap usage, and switch to a cheaper model for routine steps? One of the more useful threads I've seen on this is a Reddit post titled "What's the most an AI agent has ever quietly cost you?". The title alone is the checklist item.
- Portability. If the platform changes pricing or retires a model, can you move your instructions and files elsewhere? Instructions written in plain text and knowledge stored as ordinary files travel well. Proprietary workflow graphs don't.
Quick scorecard you can copy
Question | Good sign | Red flag |
|---|---|---|
Model choice | Many models, switchable per agent | One vendor model, no switching |
Tools and knowledge | Upload files, enable only what you need | Everything on by default |
Approvals | Agent drafts, you review | Acts on your accounts unasked |
Cost control | Visible usage, cheaper model options | Opaque credits, no per-run view |
Portability | Plain-text instructions, standard files | Locked workflow format |
Writingmate scores well on the first and last items because agents sit on top of a large model catalog (200+ models as of October 2026) and your instructions are just text. If you want help choosing among them, my earlier post on picking the right model in Writingmate is a task-first decision tree.
Step 1: Define the job in one sentence
Our example: "Every Monday, summarize what three named competitors changed on their pricing, product, and blog pages in the last seven days, and flag anything that affects our positioning."
That sentence does real work. It names a trigger (weekly), a scope (three competitors), categories (pricing, product, blog), and a judgment call (what affects us). If you can't write your agent's job this tightly, the agent can't do it either.
Write down three things before you open any tool:
- Input: the competitor names, URLs, and whatever notes you paste each Monday.
- Output: a one-page brief with fixed headings.
- Done means: every claim has a source link, and unknowns are labeled "not confirmed."
Step 2: Create the agent and write the instructions
In Writingmate, open Agents in the sidebar and choose Add Agent. According to the custom agents documentation, the required fields are name, description, instructions, and a model. Optional fields include an icon, temperature (0 for precise, 1 for creative), and context length. Agents are private by default.
The instructions field is where most of the quality comes from. The docs say good instructions cover role, audience, output format, and tool usage rules. Here's the structure I use, with the competitor agent filled in:
- Role: "You are a competitive analyst for a small software team."
- Audience: "Your reader is a founder who has five minutes on Monday morning."
- Output format: "Return a brief with these headings: Pricing changes, Product changes, Content and messaging, What this means for us, Open questions."
- Rules: "Only report changes you can support with a source link. If you can't confirm something, write 'not confirmed.' Never guess a price."
- Tool rules: "Use web search only for the competitors listed in the knowledge file."
One thing the docs flag that's easy to forget: don't paste credentials or API keys into instructions, and test with non-sensitive data first.
Step 3: Attach knowledge, not everything you own
Knowledge files give the agent reference material. Writingmate accepts files up to 10 MB each in formats such as PDF, DOCX, CSV, JSON, and Markdown. The docs are explicit that these are reference material, not permanent memory, so treat them as a library the agent consults.
For the competitor agent I'd upload just three files:
- A competitors.md file: names, URLs, and one line on why each matters.
- A positioning.md file: your own pricing, audience, and the three claims you want to protect.
- A last-week.md file: last week's brief, so the agent can say what's actually new.
Resist the urge to upload your whole drive. More files means more chances the agent pulls the wrong one, and it makes failures harder to diagnose.
Step 4: Pick a model per task
The docs say you select one model per agent, so the practical trick is to split the job. Here's how I'd think about a research agent:
Task | What matters | Model type to try |
|---|---|---|
Gathering and reading pages | Web access, long context | A search-capable or long-context model |
Spotting changes vs last week | Careful comparison | A strong reasoning model |
Writing the brief | Tone, format adherence | A faster, cheaper writing model |
You can build this as one agent per stage (a "researcher" and a "brief writer"), then paste the researcher's output into the writer. It's a manual handoff, but it keeps each agent simple and lets you swap models independently. Browse the model catalog to compare options, and check pricing for which models your plan includes, since the docs note model availability depends on your plan.
For temperature, I'd set 0 or close to it. A competitor brief should be repeatable, not creative.
Step 5: Enable only the tools this agent needs
The docs list Core Integrations, Canvas tools, Image Generation, and MCP server tools, and recommend enabling only what's necessary. For this agent, that's web access and nothing else. No image generation, no extra MCP servers. Add conversation starters like "Run this week's competitor brief" so teammates can use it without reading the instructions.
Step 6: Test with edge cases, then iterate
This is the step most people skip, and it's the one that separates a toy from a tool. Build a small test set before you trust the agent. One X post from Alexey Grigorev puts the process well:
"Start by testing your agent manually. Ask questions, review answers. But make sure you're logging everything. After 10-15 sessions, label each record as good or bad." — @Al_Grigor on X
For our competitor agent, my test set would be:
- Nothing changed. Does it say so, or does it invent news to fill the template?
- A competitor with no public pricing. Does it write "not confirmed" or guess a number?
- A page that won't load. Does it report the failure or quietly skip it?
- A misleading headline. Does it separate marketing claims from verifiable changes?
- A competitor you didn't list. Does it stay in scope?
Run each case, mark pass or fail, and change one thing at a time: an instruction line, a file, or the model. If you change three things and the result improves, you won't know why. After a few rounds, save your passing cases in a note. You'll reuse them every time you switch models.
Agentic AI vs generative AI, in plain terms
People ask "what is agentic AI vs generative AI" as if they're rival technologies. They're not. Generative AI produces content when you prompt it: a draft, a summary, an image. Agentic AI uses that same kind of model, but wraps it in a goal, instructions, tools, and a loop, so it can take several steps toward an outcome instead of answering once.
Generative AI | Agentic AI | |
|---|---|---|
Unit of work | One prompt, one response | A goal broken into steps |
Context | What you paste in | Standing instructions plus knowledge files |
Tools | Usually none | Search, files, connected apps |
Your role | Prompt every time | Set up once, review outputs |
Our competitor agent is a good example. Asking a chatbot "what did Competitor X change?" is generative. A saved agent with a role, rules, reference files, and web access that produces the same brief every Monday is agentic, even though it's still one model underneath.
Three failure modes that sink first agents
1. Vague instructions
"Be a helpful research assistant" is not an instruction. It leaves format, scope, and honesty rules up to chance. Fix: name the role, reader, output headings, and what to do when information is missing.
2. Too many tools
Every extra tool is another way for the agent to wander off. A competitor brief doesn't need image generation or ten connected apps. Fix: start with the minimum, add one tool at a time, and rerun your test set after each addition.
3. No test set
If you only try the happy path, you'll find the failures in front of your boss. Fix: five to fifteen saved cases, including at least a few that are designed to break the agent.
A fourth, quieter one is running every step on the most expensive model. Check usage after your first few runs and move routine steps like formatting to a cheaper option.
What I'd do this week
Don't start with a grand multi-agent system. Pick one repeatable job, write the one-sentence definition, and build it using the custom agent setup steps. Upload two or three files, enable one tool, set a low temperature, and run your edge-case list. Then try a second model on the same list and compare. That comparison is the cheapest way to learn what a platform's model choice is actually worth to you.
If the agent saves you even thirty minutes every Monday, you've got your answer on whether the platform is the right one. If it doesn't, your test set will tell you which part to fix.
See you in the next one!
Artem
Frequently Asked Questions
Sources
Written by
Artem Vysotsky
Ex-Staff Engineer at Meta. Building the technical foundation to make AI accessible to everyone.
Reviewed by
Sergey Vysotsky
Ex-Chief Editor / PM at Mosaic. Passionate about making AI accessible and affordable for everyone.
