Claude Code headless: how claude -p becomes an AI agent orchestrated by your code
One command. 11 complete sites. No fancy-named "team of agents" chatting with each other.
Claude Code headless is Claude Code without the interface: claude -p takes a prompt, runs the entire agent loop and returns the result on stdout, like any other process. That's what ran on my machine while I did nothing: a php artisan command that queues jobs, and each job calls that process. It finished with 78% of the subscription session used up. Billed per API token, it would have cost $528.
The trick isn't the prompt. It's treating the AI assistant as a process: instruction in, result out, and your code decides the order of things. In this post I open up the pipeline from the video: how claude -p becomes an agent orchestrated by PHP, where the math works out, and how to plug in Higgsfield as an asset generation step without letting the model decide how much to spend.
TL;DR
- What it is: Claude Code running without an interface (
claude -p), called as a subprocess by a Laravel application that controls phases, queues and budget. - Stack/Models: Laravel (artisan + queues), Claude Opus 5.5 via the Claude Code CLI, Higgsfield CLI + skills for image and video.
- Cost/Access: runs inside the Claude subscription (Max, $200) as long as you use the normal login; with
--bareor in a product for third parties, it's an API key. Higgsfield uses your plan's credits, no API key. - Useful link: headless mode documentation and the full video on YouTube.
What Claude Code headless is (and why claude -p is the Agent SDK through the terminal door)
The interactive assistant is what you use all day: open the terminal, chat, approve a tool, read the diff.
Headless is the same binary with nobody on the other end. You pass the prompt on the command line, it runs the whole loop (plans, calls a tool, reads the result, repeats) and returns the answer on stdout. End of process.
claude -p "Analise o segmento de mercado da empresa em empresa.json e escreva pesquisa.md" \
--allowedTools "Read,Write,WebSearch" \
--output-format json
Anthropic describes this literally as the Agent SDK in CLI form: "the same tools, agent loop, and context management that power Claude Code", available "as a CLI for scripts and CI/CD" (docs). And the SDK overview itself says that to use the same loop from a language other than Python or TypeScript, the way to go is to run the CLI as a subprocess with -p and --output-format json.
In other words: PHP, Bash, Go, whatever. You don't need the SDK package to have an agent. You need an exec and a JSON parser.
That changes the system design. Instead of one giant agent with 40 tools trying to remember everything, you write a deterministic pipeline in your own code and let the model solve only the probabilistic piece of each step. That's the premise of loop engineering: iterate in small steps until you hit the expected result, with your code judging when it's done.
The same reasoning applies to the visual part of the pipeline. I could spend weeks improving my image prompt for Nano Banana. Instead, I plugged in a harness that already specializes in this: Higgsfield, which aggregates the image and video models and exposes everything through a CLI and skills that Claude already knows how to use. More on it below.
The pipeline: Laravel orchestrates, Claude executes
The input is a company registry that already exists in the system. The command takes the IDs and queues one chain per company:
// app/Console/Commands/BuildSites.php
public function handle(): int
{
$ids = $this->argument('ids');
foreach ($ids as $id) {
Bus::chain([
new ResearchSegment($id),
new ResearchCompany($id),
new WriteCopy($id),
new DesignSystem($id),
new GenerateAssets($id), // etapa nova, ver abaixo
new BuildInPhases($id),
])->onQueue('sites')->dispatch();
}
return self::SUCCESS;
}
Ten companies, ten chains, ten workers. Inside each chain the phases are synchronous on purpose: the copy uses the research, the design uses the copy, the build uses all three. There's no parallelism here because there's no independence here.
Each job is a claude -p with its own prompt, its own tools and a working directory per company:
// app/Jobs/RunClaudeStep.php (base dos jobs acima)
protected function runClaude(string $prompt, array $tools, string $workdir): array
{
$result = Process::path($workdir)
->timeout(1800)
->run([
'claude', '-p', $prompt,
'--allowedTools', implode(',', $tools),
'--output-format', 'json',
'--max-turns', '40',
]);
if (! $result->successful()) {
throw new RuntimeException($result->errorOutput());
}
$json = json_decode($result->output(), true, flags: JSON_THROW_ON_ERROR);
$this->logCost($json['total_cost_usd'] ?? null, $json['session_id'] ?? null);
return $json;
}
With --output-format json the result comes back with session_id, total_cost_usd and the answer in result. The session_id lets you --resume in a later phase if you want to keep context; I prefer not to. Each phase starts clean and receives the previous artifacts as files.
What each phase does:
- Segment. What exists in that niche's market, what customers complain about, what competitors promise.
- Company. Intensive research: current site, Google reviews, social media.
- Copy. A specific prompt about headline, proof, offer. Uses the two previous files as input.
- Design. References from award-winning sites via skills (like Awesome Design MD and Awwwards) to get rid of that AI-made page look.
- Phased build. This is the part that saves the most session.
The build isn't a single claude -p. It's an implementation loop: the design becomes a plan in small phases, and each phase runs in a new process.
foreach ($plan['phases'] as $i => $phase) {
$this->runClaude(
prompt: "Implemente a fase {$i}: {$phase['goal']}. Leia PLAN.md e DESIGN.md. Não avance para a próxima fase.",
tools: ['Read', 'Write', 'Edit', 'Bash(npm *)'],
workdir: $workdir,
);
}
A small context window per phase. Opus 5.5 works with room to spare, doesn't hallucinate a component that doesn't exist, doesn't "take the opportunity" to refactor what nobody asked for. And if a phase breaks, you rerun one phase, not the whole site.
Where the math works out: subscription vs API
The numbers from the video: 11 sites, 78% of the session of a $200 Max subscription, $528 if it were API. One command paid for the subscription two and a half times over.
This works because claude -p, with no extra flags, uses the same login as your account. The tokens come out of the subscription quota.
Now the two asterisks the video didn't have time to get into:
1. --bare doesn't use the subscription. The flag recommended for CI and scripts skips hooks, skills, MCP and CLAUDE.md, and "doesn't use your subscription login": it requires ANTHROPIC_API_KEY. The documentation warns that --bare "will become the default for -p in a future release" (docs). If that happens, the subscription savings will depend on you not using the default.
2. A product for third parties means an API key. Anthropic is explicit: "Anthropic does not allow third party developers to offer claude.ai login or rate limits for their products, including agents built on the Claude Agent SDK" (overview). An internal pipeline, on your machine, generating sites for your own registry: subscription. A SaaS where the customer clicks and your Claude generates: API.
If you've read around here that claude -p was going to die, here's the update: it didn't die, it became the official entry point to the SDK for anyone not on Python or TypeScript. What's changing is the authentication default.
A tutorial shows you the way — in the Clã you build alongside us. A live class every week, real AI Engineering projects, next to people already in production.
Join the ClãThe upgrade: Higgsfield as a ready-made asset harness
The first version of the pipeline generated images by calling Gemini Nano Banana directly. It wasn't bad. But to improve it I'd have to build a harness: a prompt per asset type, style references, video for the hero, aspect ratio control. Weeks.
The same argument that made me use the Claude Code harness instead of reimplementing the loop on the Agent SDK applies here: using a ready-made, specialized harness gets better results with far less code. Higgsfield aggregates image and video models (Seedance, Kling, Veo, Nano Banana Pro, Soul) and hands them to the agent through a CLI, with skills, or through MCP.
Their official recommendation: MCP for a chat agent, CLI for a coding agent (help center). Since we're going to call this from inside a claude -p, it's the CLI.
npm install -g @higgsfield/cli
higgsfield auth login # OAuth, sem API key; usa os créditos do plano
npx skills add higgsfield-ai/skills
Done. From there, the Claude inside the pipeline has the generate, soul and product-photoshoot skills and knows how to run higgsfield generate create seedance_2_5 --prompt ... --wait.
The clip below was generated exactly this way, by the same session that wrote the post: a higgsfield generate create with Seedance 2.5, describing the pipeline in this article.
The new pipeline step comes in after the design:
// app/Jobs/GenerateAssets.php
public function handle(): void
{
$this->runClaude(
prompt: <<<PROMPT
Leia DESIGN.md e liste cada asset visual necessário (hero, produto, seções).
Para cada um, gere com a skill higgsfield-generate usando as fotos em ./fotos como referência.
Salve em ./assets e registre o job_id de cada geração em ASSETS.json.
PROMPT,
tools: ['Read', 'Write', 'Bash(higgsfield *)', 'Skill'],
workdir: $this->workdir(),
);
}
The agent decides what to generate. Laravel decides how much it can spend.
That's done with a PreToolUse hook that checks the per-site budget before allowing any higgsfield generate:
#!/usr/bin/env bash
# .claude/hooks/budget-gate.sh — recebe o tool call em stdin
input=$(cat)
cmd=$(echo "$input" | jq -r '.tool_input.command // empty')
case "$cmd" in
higgsfield\ generate*)
php artisan assets:budget-check "$SITE_ID" || {
echo "Orçamento de créditos deste site esgotado." >&2
exit 2 # bloqueia a ferramenta, o Claude recebe o motivo
}
;;
esac
Blow the budget and the tool doesn't run. It's not "remember not to spend more than X credits" in the prompt, which the model forgets by the third iteration. It's a deterministic tool controlling a probabilistic tool.
The result in the case of PimentArt, my dad's artisanal hot pepper company: three bad phone photos turned into a hero with the product placed in a scene and a short video playing behind it. The third section, which used to be an empty table, got an image explaining the three heat levels. All of it came out of the pipeline, without me opening an editor.
Limitations and things to watch out for
-p without --bare loads everything in the folder. Hooks from .claude/settings.json, servers from .mcp.json, CLAUDE.md. And it shows no trust dialog or server approval prompt. Running claude -p in a repository you don't know means executing whatever that repository configured. In your own pipeline, fine. On third-party input, no.
The session runs out. 78% of a session with 11 sites. If you queue 30, the worker will hit the rate limit halfway through and the job will fail. Treat 429 as a retry with backoff, not as a fatal error, and don't count on the session for 24/7 use.
A bad prompt generates a church. I asked for "map from above, zoom in, enter a house, person eats a coxinha with hot pepper" and the first video started in a church. The fix cost far fewer credits than generating from scratch, but it did cost. A vague description on a video model turns into burned credits.
Budget by prompt doesn't hold. Worth repeating: the only spending control that survived was the hook. An instruction in the system prompt is a suggestion.
Customer data in the prompt. Each phase receives research about the company. If that includes personal data from end customers, mask it first. The headless process has nobody watching what got sent.
Quick FAQ
Do I need the Agent SDK in Python or TypeScript to do this?
No. The packages give you approval callbacks, typed messages and streaming as objects. If all you need is prompt in, JSON out, claude -p --output-format json covers it. The documentation points to the CLI as a subprocess precisely for other languages.
Does it work with Codex or another CLI?
The orchestration does: it's Process::run with a different binary. What changes is the output format and the permission flags. The idea of small phases, artifacts in files and budget in code is model-independent.
How do I control Higgsfield credit spending?
In code, not in the prompt. A PreToolUse hook that blocks higgsfield generate when the site's budget has been blown. The CLI uses the plan credits of the logged-in account, so the real ceiling is your plan; higgsfield account status shows the balance.
Can I sell this as a product to customers? With an API key, yes. With a subscription login, Anthropic doesn't allow offering it to third parties. The pipeline stays the same; you swap the authentication and the cost line.
Conclusion
What ran here wasn't an "AI team". It was an artisan command, a queue and a binary called in a loop with a small context per phase. Claude solves the probabilistic piece of each step; PHP decides the order, keeps the artifacts and holds the budget.
When asset generation fell short, the answer wasn't to write more prompt: it was to plug in a ready-made harness and put credit control in a hook. It's the same principle on both sides.
The natural next step for this pipeline is closing the loop with evaluation: a review claude -p that scores each section and sends back the phase that failed. If you want the full mental map of what needs to exist around the model before that, start with the 5 pieces of the agent harness.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã