The 8 most popular skill packs for building software with a coding agent, each one run on real code, with what it caught and who it's for
A coding agent with no method writes code fast and builds the wrong thing just as fast. These skill packs give it a method: ask before building, test first, review before merging, check the result in a browser.
We ranked the 8 most popular packs by GitHub stars (October 2026) and ran them in Claude Code on real code. Below is what each one caught.
Superpowers stops the agent from jumping into code. Every task goes through the same steps a good engineer follows:
Brainstorm : ask why, read the code, agree on a design.
Plan : break it into small tasks with tests.
Build : one task at a time, test first, with a review between tasks.
You don't call any of this. It runs by itself.
We asked it to add team invites to a small invoicing API, the way a founder would phrase it. Before writing a line, it found a hole that would have shipped with the feature: any user could make themselves the team owner through their profile settings. It said invites shouldn't ship until that's fixed.
Then a second agent reviewed the design and found one more gap: freelancers already knew the shared login, so removing their new accounts wouldn't lock them out. Five turns later we had:
a design doc with goals and things that must not happen
a plan of 8 tasks, each starting with a failing test
edge cases for the reviewer, like an invite to Ana@Agency.com when ana@agency.com already exists
Use it if you want the agent to follow a process on every task. Skip it if most of your work is small fixes: even a one-liner goes through brainstorming unless you tell it not to.
2. Matt Pocock's skills
The most installed pack on skills.sh, with 27M installs across its skills. Superpowers is a process, while these are separate tools. You pick one up when you need it:
Skill
What it does for you
grill-me
Interviews you about an idea until every decision is made
grill-with-docs
The same, against your codebase, and it writes down the decisions
tdd
Builds a feature one failing test at a time
diagnosing-bugs
Won't guess at a fix until it can reproduce the bug
improve-codebase-architecture
Finds the refactors that will pay off
handoff
Writes up the session so another agent can continue
grill-me is the one to try first. We gave it one line: "a free month for every paying team a customer refers". In 29 seconds it came back with 7 questions, each with a recommended answer. Question 2 caught an expensive mistake: a 200-seat customer referring a 3-seat team would get $10K of credit for $90 a month of new revenue.
tdd built a parsing function in 7 small red-to-green steps: 17 passing tests in under two minutes. It asked us to confirm the rules before writing a single test.
improve-codebase-architecture read the git history of the zustand library, focused on the files that change most, and wrote an HTML report with 5 candidates. Two were rated Strong, and one was linked to two real past bugs.
0:00 / 0:00
Use it if you want the discipline without a process you can't switch off. Start with grill-me. It works on any decision, not only code.
3. caveman
caveman makes the agent talk less: no greetings, recaps or filler, just the answer. It went viral as a way to cut your AI bill.
We measured it. We asked the same three questions about a real codebase six times each:
Condition
Output tokens
Cost at API prices
No instruction
9,367
$1.07
/caveman
8,150 (โ13%)
$1.05 (โ1%)
One line: "Answer concisely."
5,814 (โ38%)
$0.91 (โ15%)
The answers stayed correct, but a plain "Answer concisely." beat it on every measure. Cost barely moved either way, because most of an agent's bill is the code it reads, not what it writes.
Skip the skill. Put "Answer concisely" in your CLAUDE.md. If your bills are mostly input, caveman also has a proxy that compresses what the agent reads. We didn't test it.
4. Addy Osmani's agent-skills
Addy Osmani is a longtime engineering lead at Google. This pack is his way of shipping, as commands: /spec, /plan, /build, /test, /review, /ship. Each step has checks the agent can't skip. It even lists the excuses agents use to skip a step, with an answer to each one.
/review is the one to use every day. We gave it a branch with bugs we'd planted on purpose. In 52 seconds (about 36 cents of usage at API prices) it found all of them, plus two we hadn't planted:
Planted: the new CSV export could never be reached, because another route caught it first.
Planted: any user could send reminder emails to other teams' clients.
Not planted: invoices counted as overdue from the draft date, not the send date.
Not planted: timestamps read in the wrong timezone. It ran the code in New York time to prove it.
Use it if you merge agent-written code. A one-minute review caught bugs that would otherwise have shipped.
5. archify
archify turns a codebase, or a plain description like "browser โ API โ cache โ database", into an interactive diagram. Click any box to see what feeds it and what it feeds. Export it as a picture when you need one for a doc or a pitch.
We pointed it at zustand. In under 2 minutes it traced 11 parts of the library back to their source files, highlighted the main path a state update takes, and marked the outside dependencies.
Use it when you onboard a new developer, explain your system to an investor, or inherit a codebase. It draws what's in the code, not what's deployed, so check it against reality.
6. agent-browser
Vercel's agent-browser gives the agent a real Chrome to open, click, type in, screenshot and record. The agent stops saying "should work now" and checks for itself: it signs up, clicks through checkout, and sends you a screenshot.
We asked it to find the most-installed changelog skill on skills.sh. It took 52 seconds and 10 steps, and it skipped an off-topic result on its own. agent-browser recorded the video below itself:
0:00 / 0:00
Use it as soon as you have a UI. Ask "sign up with a test email and screenshot the dashboard" after every change to the flow. Keep it away from accounts where a wrong click costs money.
7. Cloudflare security-audit
The security audit Cloudflare runs on its own code, packaged for one repo. One agent hunts for holes, and a separate one tries to prove each finding wrong. Only what survives gets reported, so you don't drown in false alarms.
We ran it on an API we'd seeded with vulnerabilities. It confirmed 11 findings with file and line, including every serious one we planted:
reading another team's invoices by changing an ID
making yourself team owner
a payment webhook anyone could fake to mark invoices paid
a database injection in search
It also found one we didn't plant: a stolen login token stayed valid after a password reset.
Then Claude's safety filter blocked the final write-up, twice. The findings were there, but we never got the polished report.
Use it before launch, and run it more than once: Cloudflare found that a single run catches about half of what repeated runs find. For a lighter first pass on a fast-built app, see vibe-security in top skills for SaaS .
8. Trail of Bits skills
Trail of Bits is one of the best-known security audit firms. Their pack has 93 security skills, and most of them are for smart contracts and low-level code. For a startup web app, three matter:
differential-review: a security review of your latest change.
insecure-defaults: finds settings that fail open, like auth turned off or a debug mode left on.
supply-chain-risk-auditor: flags risky or abandoned npm and Python packages.
We didn't run it for this article. Use it for a focused check on a sensitive change, after the broad Cloudflare audit.
Which to install first
grill-me and tdd from Matt Pocock's skills, if you install only two. They fix the most common failures: building the wrong thing, and code that doesn't work.
Superpowers instead, if you want the process enforced on every task. Pick one or the other, not both.
Addy Osmani's /review before every merge.
agent-browser as soon as you have a UI.
Skip caveman's skill. One line in CLAUDE.md does more.
Ship faster than your competition
Focus on customers and sales while we handle product delivery. Hire a dedicated AI maker
or a whole product team.