Agentic Engineering
Open skills near APIs, scrapers that explore, visual plans, screenshot QA, and ugly files that remember what the model forgets.
If someone only reads one block, this should still be enough to start building.
I’ve watched a lot of teams buy an “agent framework,” paste a giant system prompt, and wonder why the bot still invents endpoints that don’t exist. The failure mode isn’t the model. It’s that the agent never got tools that look like the real world, or files that remember last week’s pain.
Inside Masse we started calling the opposite approach Agentic Engineering — engineering shaped around agents, not around frameworks. Resourceful. Slightly chaotic. Obsessed with what actually ships. No cathedral. Just open skills, scrapers, visual plans, screenshots, and a handful of markdown files that tell the truth.
The symptoms look like a tooling problem
When an agent fails, people upgrade the model. Half the time they should have upgraded the surroundings.
| Symptom | Cause |
|---|---|
| Agent invents endpoints and field names | No skill that mirrors the real API shape |
| Same bug explained every Monday | Quirks lived in Slack, never in a file the agent reads |
| “It said it fixed it” with no proof | No screenshot or visual plan skill for QA |
| Scrape jobs hallucinate page structure | Gathering without an explore step first |
Treat the agent like a junior hire: when it guesses, you starved it of tools and notes.
Skills should feel like APIs
Here’s the rule I keep repeating: open skills beat closed frameworks when those skills sit close to the APIs and CLIs the work already uses. If Google Search Console has a query shape, the skill should expose that shape. If Firecrawl returns markdown and links, the skill should return markdown and links — not a proprietary “DocumentObject.”
Closed stacks feel productive for a week. Then you need one weird header, one auth quirk, one pagination edge case, and you’re fighting the framework instead of teaching the agent. Open skills you can fork. You can paste the failing request into the skill file. You can version them next to the repo the agent works in.
If a skill can’t be explained as “thin wrapper around X,” it’s already too clever.
The stack that survives contact
The perfect agentic engineering stack is boring on purpose. Five layers, each one a skill or a file the agent can touch.
Scrape without explore is guessing. Explore without gather is tourism. Plans without screenshots are fan fiction.
flowchart LR
A[Scrape] --> B[Explore]
B --> C[Gather]
C --> D[Visual plan]
D --> E[Screenshot QA]
E --> F[Update findings]
F --> A
G[Quirks + templates + goals] -.-> C
G -.-> D
The loop matters more than any single tool. Findings feed the next scrape.
Scrape, then explore, then gather
Firecrawl (or anything in that family) is the doorway. You need raw pages, maps of URLs, structured extracts when the schema is clear. But scraping alone is how agents invent a site map from three lucky hits.
So you split the work into two skills on purpose:
- Explore — what exists around the target? Routes, folders, sibling pages, related APIs.
- Gather — given a specific task, pull only the evidence that moves that task forward.
Explore answers “what neighborhood am I in?” Gather answers “what do I need for this ticket?” Mix them and the agent either over-fetches or under-reads. Keep them separate and you can reuse explore output across three gather passes without paying for the tour again.
Illustrative, not a lab study — but it matches what we see when the same model gets better surroundings.
Documentation skills: plans and screenshots
Agents that can’t show their work will lie politely. Two skills fix most of that:
A visual plan skill forces the agent to draw the approach before it edits — layout, sequence, risks, what “done” looks like. You review the plan the same way you review a PR description, except it’s visual enough that gaps jump out.
A screenshot skill is for QA and reporting. Before and after. Broken state and fixed state. The agent attaches proof instead of narrating confidence. This is also how you keep humans in the loop without making them re-run the whole session.
Together they are the documentation layer of Agentic Engineering: not a Confluence page nobody opens, but artifacts produced in the same run as the work.
The files that keep the agent honest
Tools move tokens. Files hold judgment. The minimum set we keep next to agent work:
quirks.md
Weird edges that already bit us — auth dances, null fields, pagination surprises.
templates/
Known-good shapes for tickets, reports, and PRs so the agent starts from something that already works.
findings.md
Running notes from this week so the next session doesn’t rediscover the same dead end.
goals.md
What “done” looks like — specific enough that a screenshot can falsify it.
Name them whatever you want. Keep the jobs. Instructions without these files are just hope.
Quirks are the API oddities, the auth dance, the field that’s sometimes null on Tuesdays. Templates are the known-good shapes for tickets, reports, PRs. Findings are running notes so the next agent session doesn’t rediscover the same dead end. Goals are the scoreboard — specific enough that a screenshot can falsify them.
BEFORE YOU TOUCH CODE:
1. Read goals.md — what does done look like?
2. Skim quirks.md for the surface you are about to call.
3. Explore, then gather. Do not scrape blind.
4. Draft a visual plan. Wait for a human if risk is high.
5. Ship, then screenshot. Append findings.md.
6. Never invent an API field. Open the skill or the docs.
Paste this into the agent brief. Short enough to survive a context trim.
Extra tip: docs, sheets, and the company wiki
Agents waste hours regenerating tables that should have been a sheet
edit. Use a Workspace CLI when you need fast sheet and doc generation
or surgical edits — in our setup that’s
gws. The agent should call it the way a
human would: create a tab, append a row, fix one cell, move on.
Same idea for docs: edit the section, don’t regenerate the manuscript.
And wire in your company wiki as a first-class source. Agents that only see the ticket will invent policy. Agents that can search the wiki for how your team already solved auth, naming, or QA will sound like they work here. Point the gather skill at the wiki the same way you point it at the codebase.
What this looks like on a real ticket
Say the ticket is “fix the client report that loses three rows
on export.” Agentic Engineering doesn’t start with a
40-line prompt. It starts with explore on the report route and the
export handler, gather on the failing fixture, a skim of
quirks.md for known CSV edges, a visual
plan of the fix, the code change, then a screenshot of the export
preview with the three rows present. Findings get a line. Goals get
checked off. Next week’s agent inherits the lesson.
That’s the whole pitch. Frameworks age. Open skills and ugly files compound.
If you only do one thing after reading this: pick a single API your
agents already abuse, wrap it in an open skill that matches the real
request shape, and create an empty
quirks.md next to it. The rest of the
stack gets easier once those two exist.
Next: keep that memory in a wiki the agent can actually update