An AI agent is an AI that takes actions toward a goal instead of just answering. You give it a task; it plans, uses tools — reading and editing files, running commands, checking results — and keeps going until the task is done or it needs you.
If you've used the copy-paste workflow — draft in one model, paste into another to review, paste again to polish — you've already done an agent's job by hand. An agent does that loop for you, inside your actual project, and shows you the evidence at the end.
Chat vs. Agent: What Actually Changes
The models are often the same. What changes is where the work happens and who does the busywork.
| Question | Chat / Custom GPT | Agent |
|---|---|---|
| Where does the work happen? | In a chat window | In your project files |
| Who copies and pastes? | You | Nobody — the agent edits directly |
| Can it run tests and commands? | Rarely, in a sandbox | Yes, in your project (with your approval) |
| Where do standing rules live? | The GPT's instructions | A context file in your project |
| Worst case if it's wrong? | A bad answer you can ignore | A bad change to real files — so guardrails matter |
That last row is the whole story. Agents are more useful because they can change things, which is also why the rest of this guide spends as much time on checking as on building.
The Kinds of Agents You'll Meet
Product names change every few months; these roles don't. Pick by role, then by how well the tool lets you approve and review its work.
🛠️ Agent mode in your editor
The gentlest start. Inside VS Code or a similar editor, you describe a task and watch the agent edit files and run commands, approving as it goes.
Examples: Claude Code or Codex extensions for VS Code, Copilot agent mode, Cursor
🖥️ Terminal coding agents
Work across a whole project from the command line. Strong at multi-file changes, running test suites, and following project rules.
Examples: Claude Code, Codex CLI, Gemini CLI
☁️ Cloud & background agents
Assign a task — often from an issue — and the agent works in its own sandbox, then hands back a pull request for you to review.
Examples: Codex cloud, Claude Code on the web, Copilot coding agent
⚙️ Orchestrated agents
Several agents run in a fixed order — one builds, another reviews, a script checks — with a gate at every hand-off. This is deterministic orchestration.
Examples: builder + reviewer pipelines, scheduled runs
Briefing an Agent (Not Prompting It)
A prompt asks for an answer. A brief hands over a job. The difference is that a brief says what "done" means and where the agent has to stop and ask. A short, specific brief beats a long, vague one every time.
Goal: [One sentence. e.g. Add a "Holiday Hours" banner to the homepage.] Context: [What the project is. e.g. Static site, plain HTML/CSS. See AGENTS.md for project rules.] Constraints: - [e.g. Match the existing colors and fonts.] - [e.g. Don't change any other page.] - Work on a new branch. Don't push. Done when: - [e.g. The banner shows correctly on desktop and mobile.] - [e.g. No broken links or console errors.] - You show me what changed and how you checked it. Ask me before: deleting files, installing anything, or pushing.
Notice the last two sections. "Done when" turns a vague request into something the agent can check itself against. "Ask me before" draws the line between what it may do alone and what needs you.
Context Files: Custom Instructions for Your Project
A Custom GPT has system instructions. An agent has a context file: a plain text file in your project that the agent reads at the start of every session. Many agents look for AGENTS.md; some use their own name, like CLAUDE.md. Check your tool's docs for which file it reads.
Write it once and every session starts informed. Keep it short — rules, not essays:
# Project notes for AI agents ## What this is A static website for a small bakery. Plain HTML + CSS, no build step. ## Rules - Keep the existing look: colors and fonts live in css/styles.css. - Every page needs a title, a meta description, and exactly one <h1>. - Never edit anything in /archive. - Ask before deleting files or adding any new dependency. ## How to check your work - Preview locally: python3 -m http.server 8000 - Click every link and check every image on the pages you changed. ## Done means Changes are on a branch, checks pass, and you've listed what changed and how you verified it.
MCP: How Agents Plug Into Your Tools
MCP, the Model Context Protocol, is an open standard for connecting AI apps and agents to tools and data. An MCP server exposes one capability — your code host, a database, a docs site, a browser — and any MCP-capable agent can use it.
Think of it as the agent-era version of a Custom GPT's "Actions," except it works across tools from different vendors. Connect your issue tracker once, and any agent that speaks MCP can read your issues.
Guardrails That Keep You in Control
Agents fail in a predictable way: they're confident, fast, and sometimes wrong about something you didn't think to tell them. Guardrails turn that from a disaster into a diff you reject.
Before the agent starts
- Work on a branch or a copy — never directly on production
- Commit or back up first, so any change can be undone
- Keep approval prompts on for commands, deletions, and pushes
- Don't hand it secrets or production credentials it doesn't need
- Write the brief with a clear "done when"
Before anything ships
- Read the summary and the actual changes (the diff)
- Ask for evidence: test output, screenshots, live checks
- Watch for scope creep — changes you didn't ask for
- Test on a preview branch before main
- You approve the push, not the agent
Verify: Make the Agent Prove It
"Done" should mean verified, not "the agent says so." The habit that matters most is asking for evidence you can read yourself:
- Before and after: what did the live site, page, or report look like before the change, and after?
- Test output: the actual results, not a summary that they passed.
- The diff: exactly which lines changed, in which files.
- What it didn't do: steps it skipped, checks it couldn't run, anything it's unsure of.
A trustworthy agent will tell you when it couldn't verify something. Treat an agent that never admits uncertainty with suspicion.
A Real Example: Maintaining This Site with an Agent
This guide was written alongside a real maintenance job on this site, done with a terminal coding agent (Claude Code) and a human approving each stage. The task: move two small files (robots.txt and security.txt) off a shared Cloudflare Worker and into the site itself, without touching anything else.
-
Review before touching anything
The agent found the working folder held three dated copies of the site, checked which one matched GitHub, and used that one — instead of editing the first copy it found.
-
Capture a baseline
It recorded what the live site served before the change, so "after" would mean something.
-
Change, then stop at the irreversible step
It made the four requested changes, showed the diff, committed locally — and stopped before pushing to production to ask for a yes.
-
"Check your work" found real problems
Asked to double-check, it found the project README described an outdated setup — and that one of its own sentences claimed something it couldn't actually verify. Both were corrected.
-
Prove it live
After approval it pushed, waited for the deploy, and ran the exact verification checks from the original brief, confirming the site's security headers were unchanged.
From One Agent to Orchestration
Once a single agent is reliable on a task, the next step is deterministic orchestration: the same agents, in the same order, with the same checks, every run. One agent builds; a second, ideally on a different model, reviews; a script or test suite decides whether it passes. It's the multi-model review workflow you may already use — with the copy-paste automated and a real gate at every hand-off.
Ready to Try Your First Agent Task?
Pick something small and reversible — a new section on one page, a fixed broken link — and ship it through a preview branch.
GitHub Preview Workflow → See the Learning Path →Frequently Asked Questions
What is the difference between a Custom GPT and an AI agent?
A Custom GPT answers inside a chat window; you copy its output and do the work yourself. An agent works directly in your project: it reads and edits files, runs commands and tests, and reports back what it did. Both use standing instructions, knowledge, and tools, but the agent keeps them in your project and acts on them.
Do I need to know how to code to use an AI agent?
No, but you do need to know what you want and how to check it. The skill shifts from writing code to writing a clear brief, reading what the agent reports, and verifying the result. Agent mode inside an editor like VS Code is the gentlest place to start.
Is it safe to let an agent edit my files?
It can be, with guardrails. Work on a branch or a copy, keep approval prompts on for commands and deletions, never give the agent secrets or production credentials it does not need, and review every change before it ships. You should be the one who approves anything irreversible, such as a push to production.
What is MCP?
MCP, the Model Context Protocol, is an open standard for connecting AI apps and agents to tools and data. An MCP server exposes something like your code host, a database, or a docs site, and any MCP-capable agent can use it. Install servers only from sources you trust and give them the narrowest access that works.
Which AI agent should I start with?
For most people, Claude Code or Codex. Both run as a VS Code extension, which is the gentlest start, and in the terminal once you are comfortable. Many builders use one to build and the other to review. Product names change quickly, so pick by role and by how well the tool lets you approve and review its work.