Trust is earned, not given

A different perspective

2026-06-30 · AI

Autonomous Execution Loops in Modern AI Agents

Autonomous Execution Loops in Modern AI Agents: Architecture, Implementation, and Evaluation of Claude Code's /goal Command

The evolution of artificial intelligence in software engineering has rapidly progressed from passive code completion to interactive conversational agents, and most recently to autonomous, agentic coding environments. In early iterations of developer tools, AI models operated strictly within single-turn context interactions: a developer submitted a prompt, the model generated a response or suggested code edits, and the execution loop terminated immediately. While effective for localized functions and minor refactoring, this synchronous single-turn model imposed significant operational friction on complex engineering workflows. Tasks requiring multi-file modifications, iterative test-driven development (TDD), long-running debugging cycles, or continuous linting required human developers to act as constant manual orchestrators, prompting the AI assistant repeatedly after every individual tool execution. To bridge this gap between interactive assistance and autonomous software synthesis, modern agentic environments have introduced persistent completion hooks. Prominent among these implementations is Anthropic's `/goal` command within Claude Code, a feature introduced in version 2.1.139 that fundamentally transforms the developer-agent interaction paradigm from conversational prompting to goal-driven autonomous loop execution.

The `/goal` command allows developers to specify a persistent, high-level completion condition—such as 'make all unit tests in test/auth pass and clean up all linter warnings'—and instructs the AI agent to iteratively execute tool actions across multiple turns until the condition is verifiably met. Unlike naive infinite loops or simple auto-continuation scripts, the `/goal` command introduces a sophisticated, dual-model architecture that decouples action execution from evaluation. While a primary coding model (such as Claude 3.5 Sonnet) executes tool calls and modifies the codebase, a second, lightweight evaluation model (such as Claude 3.5 Haiku) acts as an independent adversarial auditor, evaluating conversation transcripts and state changes after every turn to verify objective completion. This paper presents an in-depth analysis of the `/goal` command in Claude Code, examining its operational mechanics, system architecture, underlying hook mechanisms, session lifecycle management, and security boundaries. Furthermore, this paper provides a concrete TypeScript implementation demonstrating how developers can engineer an enterprise-grade autonomous goal loop with dual-model verification, state persistence, error handling, and token budget safety controls.

To understand the technical significance of the `/goal` command, one must contrast traditional command-line AI workflows with autonomous goal-oriented loops. In a conventional CLI session, an AI assistant executes an action (e.g., editing a source file or invoking a compiler) and yields control back to the user terminal. If a compiler error occurs, the user must manually copy the error log, formulate a new prompt, and trigger the next turn. This approach creates an operational bottleneck in long-running tasks, such as resolving legacy technical debt, fixing complex test suites, or conducting comprehensive repository refactoring.

The `/goal` command eliminates this manual polling cycle by converting a prompt into a session-scoped, continuous execution directive. When a developer executes `/goal [completion condition]`, Claude Code initializes a persistent evaluation loop. During each turn, the primary agent receives the goal context, inspects the project environment using available Model Context Protocol (MCP) tools—such as file system readers, bash execution terminals, or static analysis linters—and executes necessary code modifications. Once the primary agent completes its turn and yields control, the engine automatically intercepts the execution flow before returning control to the human user. It dispatches the cumulative conversation transcript, tool execution results, and original goal criteria to an isolated evaluator agent. If the evaluator determines that the completion criteria remain unsatisfied, the engine automatically re-prompts the primary agent with the evaluator's feedback, initiating the subsequent turn without requiring human intervention.

The user interface provides real-time visibility into this autonomous lifecycle. An active goal status indicator (`â—Ž /goal active`) displays the running elapsed time, turn counter, and accumulated token spend. Claude Code supports subcommands for dynamic goal management, including `/goal` (without arguments to view status), `/goal pause`, `/goal resume`, and `/goal clear` (or its aliases `stop`, `reset`, or `cancel`). Furthermore, the command functions seamlessly across interactive terminal sessions, non-interactive CI/CD scripts via `claude -p "/goal ..."`, and remote control environments, establishing a versatile foundation for modern vibe coding and automated software maintenance (Anthropic).

The core architectural innovation of the `/goal` command lies in its resolution of the 'self-judgment bias' inherent in single-agent architectures. When a single language model is tasked with both executing code and determining whether its own output fulfills a complex directive, it exhibits a strong confirmation bias. Models frequently claim that a task is finished based on plausible reasoning or visual inspection of uncompiled code, prematurely terminating execution despite existing runtime failures or incomplete implementations.

Claude Code resolves this bias by implementing a decoupled, dual-model evaluation architecture grounded in prompt-based `Stop` hooks. The primary agent (typically a highly capable model like Claude 3.5 Sonnet) operates in 'execution mode,' possessing full access to file editing, file creation, shell command execution, and code analysis tools. However, the decision to terminate the loop does not reside with Sonnet. Instead, upon turn completion, the Claude Code runtime invokes a secondary, lightweight model (defaulting to Claude 3.5 Haiku via the Claude API) acting exclusively as an evaluation judge.

This evaluator model is purposefully sandboxed: it does not possess tool-calling permissions or write access to the workspace. Its sole responsibility is to ingest the goal text, the system state, and the full transcript of recent turns, evaluating whether the completion criteria have been satisfied in reality rather than in theory. The evaluator outputs a structured JSON response containing a boolean completion flag (`is_achieved`), a confidence score, and detailed progress feedback. If `is_achieved` is `false`, the evaluator's feedback is injected into the primary agent's prompt context for the next turn. If `is_achieved` is `true`, the runtime prints a confirmation message, clears the session goal, and yields control back to the user. By delegating judgment to an independent, non-execution model, Claude Code guarantees rigorous verification, preventing early termination traps (ExplainX).

Operating an autonomous execution loop in production requires robust error handling, session persistence, and resource management. The `/goal` command implements a comprehensive state machine that manages turn execution across diverse operational edge cases, preventing runaway token consumption and infinite loops.

First, the engine distinguishes between recoverable tool errors and unrecoverable environmental failures. If a primary agent turn results in a transient tool failure (such as a temporary network timeout or a syntax error emitted by a test runner), the loop persists. The error output is fed back into the conversation context, allowing the agent to self-correct on the subsequent turn. However, if an unrecoverable error occurs—such as missing file permissions, corrupted settings, or an explicit configuration error requiring manual user intervention—the runtime safely halts execution, clears the goal state, and logs a warning: `Goal cleared after an unrecoverable error. Run /goal again to continue.` This design prevents the agent from spinning in an expensive, unrecoverable error loop.

Second, the `/goal` command integrates deeply with session persistence primitives. If a developer terminates a CLI session or if a network connection drops while a goal is active, Claude Code serializes the goal condition and its evaluation metadata into the local session transcript file. Upon resuming the session through any route (`claude --continue`, `claude --resume [id]`, or the session picker UI), the engine restores the active goal state. Crucially, while the goal condition persists across session restores, the turn counter, timer, and token-spend baselines are reset, ensuring accurate resource tracking for the resumed execution context (Anthropic).

Third, security boundaries are maintained through interaction modes. When executed in standard manual mode, Claude Code still prompts the user for permission before executing sensitive tool actions (such as arbitrary shell commands or file deletions). To achieve true unattended autonomy, developers combine `/goal` with auto-permission mode (`claude --auto-approve` or running in headless CI environments). In this configuration, permission prompts are automatically approved within turn boundaries, while the evaluator model continuously audits system outputs, striking an optimal balance between developer velocity and system safety.

To illustrate how the `/goal` architecture can be implemented from first principles, we present a complete Node.js/TypeScript implementation. The code below models the core mechanics of Claude Code's `/goal` command, featuring a session manager, primary execution agent, secondary evaluator agent, state persistence file system, and token budget limits.

Listing 1: Complete TypeScript Implementation of an Autonomous Goal Execution Loop

import { Anthropic } from '@anthropic-ai/sdk';
import * as fs from 'fs';
import * as path from 'path';

// System Interfaces
interface GoalSession {
  sessionId: string;
  condition: string;
  isActive: boolean;
  turnCount: number;
  totalTokensUsed: number;
  maxTurns: number;
  maxTokenBudget: number;
  createdAt: string;
}

interface EvaluationResult {
  isAchieved: boolean;
  isImpossible: boolean;
  feedback: string;
  confidence: number;
}

export class AutonomousGoalEngine {
  private anthropic: Anthropic;
  private session: GoalSession;
  private transcript: Anthropic.MessageParam[] = [];
  private stateFilePath: string;

  constructor(apiKey: string, sessionId: string, workspacePath: string) {
    this.anthropic = new Anthropic({ apiKey });
    this.stateFilePath = path.join(workspacePath, `.goal_session_${sessionId}.json`);
    this.session = this.loadOrCreateSession(sessionId);
  }

  // Set or replace an active persistent goal
  public async setGoal(condition: string, maxTurns = 20, maxTokenBudget = 100000): Promise<void> {
    if (!condition || condition.trim().length === 0) {
      throw new Error("Goal condition cannot be empty.");
    }
    if (condition.length > 4000) {
      throw new Error("Goal condition exceeds maximum allowed length of 4000 characters.");
    }

    this.session = {
      sessionId: this.session.sessionId,
      condition: condition.trim(),
      isActive: true,
      turnCount: 0,
      totalTokensUsed: 0,
      maxTurns,
      maxTokenBudget,
      createdAt: new Date().toISOString()
    };

    this.saveSessionState();
    console.log(`[Goal Engine] Active goal set: "${this.session.condition}"`);
    
    // Immediately trigger autonomous execution loop
    await this.runGoalLoop();
  }

  // Core Autonomous Loop Logic
  private async runGoalLoop(): Promise<void> {
    while (this.session.isActive) {
      this.session.turnCount++;
      console.log(`\n--- [Turn ${this.session.turnCount}/${this.session.maxTurns}] Executing Goal Loop ---`);

      // 1. Safety Boundary Checks
      if (this.session.turnCount > this.session.maxTurns) {
        console.warn("[Goal Engine] Maximum turn limit reached. Clearing goal.");
        this.clearGoal("Turn limit exceeded.");
        break;
      }

      if (this.session.totalTokensUsed >= this.session.maxTokenBudget) {
        console.warn("[Goal Engine] Token budget exceeded. Clearing goal.");
        this.clearGoal("Token budget depleted.");
        break;
      }

      try {
        // 2. Primary Agent Turn Execution (Claude 3.5 Sonnet)
        const agentResponse = await this.executePrimaryAgentTurn();
        this.session.totalTokensUsed += (agentResponse.usage?.input_tokens || 0) + (agentResponse.usage?.output_tokens || 0);

        // Append Primary Agent output to Transcript
        this.transcript.push({
          role: 'assistant',
          content: agentResponse.content
        });

        // 3. Secondary Evaluator Agent Turn (Claude 3.5 Haiku) - Decoupled Audit
        const evaluation = await this.evaluateGoalProgress();
        console.log(`[Evaluator] Achieved: ${evaluation.isAchieved} | Feedback: ${evaluation.feedback}`);

        if (evaluation.isAchieved) {
          console.log(`\n SUCCESS: Goal achieved in ${this.session.turnCount} turns!`);
          this.clearGoal("Goal successfully achieved.");
          break;
        }

        if (evaluation.isImpossible) {
          console.error(`\n ABORT: Evaluator judged goal impossible: ${evaluation.feedback}`);
          this.clearGoal("Goal determined to be impossible.");
          break;
        }

        // 4. Inject Evaluator Feedback into Context for Next Turn
        this.transcript.push({
          role: 'user',
          content: `[SYSTEM GOAL EVALUATOR FEEDBACK]: Condition not met yet. ${evaluation.feedback}. Please continue working toward: "${this.session.condition}"`
        });

        this.saveSessionState();

      } catch (error: any) {
        console.error(`[Goal Engine] Error during execution turn: ${error.message}`);
        if (this.isUnrecoverableError(error)) {
          console.error("Unrecoverable error encountered. Clearing goal state.");
          this.clearGoal(`Unrecoverable error: ${error.message}`);
          break;
        }
        // Recoverable error: prompt agent to resolve issue in next turn
        this.transcript.push({
          role: 'user',
          content: `[SYSTEM ERROR]: The previous turn failed with error: ${error.message}. Please fix this error and continue.`
        });
      }
    }
  }

  // Primary Coding Agent Execution (Simulated Tool Access Context)
  private async executePrimaryAgentTurn(): Promise<Anthropic.Message> {
    const systemPrompt = `You are an autonomous software engineering agent. Your objective is to fulfill the user's goal: "${this.session.condition}".
Use available tools, inspect files, edit code, and run test suites. Work systematically until complete.`;

    if (this.transcript.length === 0) {
      this.transcript.push({
        role: 'user',
        content: `Please execute all necessary steps to fulfill this goal: "${this.session.condition}"`
      });
    }

    return await this.anthropic.messages.create({
      model: 'claude-3-5-sonnet-20241022',
      max_tokens: 4096,
      system: systemPrompt,
      messages: this.transcript
    });
  }

  // Secondary Evaluator Agent (Decoupled Claude 3.5 Haiku Audit)
  private async evaluateGoalProgress(): Promise<EvaluationResult> {
    const evaluatorSystemPrompt = `You are an independent, adversarial goal evaluator.
Your role is to inspect the conversation transcript and determine if the user's condition has been FULLY satisfied in reality.
Be rigorous. Do not assume code works unless test execution output or build clean logs prove it.

Goal Condition: "${this.session.condition}"

Respond ONLY with a JSON object in this exact schema:
{
  "isAchieved": boolean,
  "isImpossible": boolean,
  "confidence": number (0.0 to 1.0),
  "feedback": "Concise summary of what remains to be done or why it is achieved"
}`;

    // Flatten transcript for evaluator inspection
    const conversationSummary = this.transcript
      .map(m => `${m.role.toUpperCase()}: ${typeof m.content === 'string' ? m.content : JSON.stringify(m.content)}`)
      .slice(-6) // Evaluator inspects recent turns context window
      .join('\n\n');

    const evalResponse = await this.anthropic.messages.create({
      model: 'claude-3-5-haiku-20241022',
      max_tokens: 1000,
      system: evaluatorSystemPrompt,
      messages: [
        {
          role: 'user',
          content: `Recent Conversation Transcript:\n${conversationSummary}\n\nHas the goal condition been met?`
        }
      ]
    });

    const textContent = evalResponse.content[0].type === 'text' ? evalResponse.content[0].text : '{}';
    try {
      const parsed = JSON.parse(textContent);
      return {
        isAchieved: Boolean(parsed.isAchieved),
        isImpossible: Boolean(parsed.isImpossible),
        confidence: Number(parsed.confidence) || 0.5,
        feedback: parsed.feedback || "No detailed feedback provided."
      };
    } catch {
      // Fallback parser if JSON formatting contained markdown blocks
      const isAchieved = textContent.includes('"isAchieved": true');
      return {
        isAchieved,
        isImpossible: false,
        confidence: 0.5,
        feedback: "Parsed from unstructured evaluator response."
      };
    }
  }

  private isUnrecoverableError(error: any): boolean {
    return error.status === 401 || error.message.includes("ENOENT") || error.message.includes("EACCES");
  }

  public clearGoal(reason: string): void {
    this.session.isActive = false;
    this.saveSessionState();
    console.log(`[Goal Engine] Goal cleared: ${reason}`);
  }

  private saveSessionState(): void {
    fs.writeFileSync(this.stateFilePath, JSON.stringify(this.session, null, 2), 'utf-8');
  }

  private loadOrCreateSession(sessionId: string): GoalSession {
    if (fs.existsSync(this.stateFilePath)) {
      try {
        const raw = fs.readFileSync(this.stateFilePath, 'utf-8');
        return JSON.parse(raw);
      } catch {
        // Fallback on corrupt file
      }
    }
    return {
      sessionId,
      condition: '',
      isActive: false,
      turnCount: 0,
      totalTokensUsed: 0,
      maxTurns: 20,
      maxTokenBudget: 100000,
      createdAt: new Date().toISOString()
    };
  }
}

The TypeScript implementation above demonstrates the four essential components of an autonomous goal-driven agent system. First, the `AutonomousGoalEngine` class maintains state persistence via local JSON serialization (`.goal_session_[id].json`), enabling seamless session recovery across CLI restarts or system failures. Second, the execution loop in `runGoalLoop()` enforces strict safety budgets, checking both turn bounds (`maxTurns`) and cumulative token consumption (`maxTokenBudget`) prior to issuing API calls. This prevents runaway billing loops in scenarios where an agent becomes stuck in an unresolvable error loop.

Third, the code highlights the crucial separation between `executePrimaryAgentTurn()` and `evaluateGoalProgress()`. The primary agent runs on `claude-3-5-sonnet-20241022`, utilizing full prompt context to formulate execution plans and edit code. In contrast, the evaluation step executes via `claude-3-5-haiku-20241022` with a specialized system prompt enforcing strict JSON output validation. This dual-model design guarantees that evaluation remains lightweight, cost-effective, and adversarial, preventing the primary agent from self-approving incomplete or broken code.

The introduction of the `/goal` command in Claude Code marks a pivotal shift in software engineering automation. By transitioning from reactive single-turn prompts to persistent, evaluated completion conditions, AI developer tools empower engineers to delegate high-level outcomes rather than step-by-step instructions. This shift enables true 'vibe coding' and autonomous maintenance—allowing developers to initiate multi-hour refactoring tasks, test-suite resolutions, or dependency migrations with a single command.

However, as autonomous loops become standard across commercial developer environments (such as Claude Code and Codex CLI), engineering teams must implement robust safety harnesses, token spend monitors, and continuous integration audits. While secondary evaluator models significantly mitigate self-judgment bias, human code review and architectural oversight remain essential for ensuring that autonomously generated code meets long-term maintainability and security standards. The `/goal` architecture establishes a scalable blueprint for human-AI collaboration, where humans define objective completion criteria, and autonomous agents work tirelessly until those standards are verifiably achieved.

Works Cited

Anthropic. "Keep Claude Working Toward a Goal - Claude Code Documentation." Anthropic Documentation, 25 Sept. 2026, code.claude.com/docs/en/goal. Accessed 27 Sept. 2026.

Anthropic. "Slash Commands and CLI Configuration in Claude Code." Anthropic Documentation, 24 Sept. 2026, code.claude.com/docs/en/commands. Accessed 27 Sept. 2026.

Daily.dev. "Goal Mode Changes Everything for AI Coding." Daily.dev Tech News, 13 May 2026, daily.dev/posts/goal-mode-changes-everything-for-ai-coding-2u794aqel. Accessed 25 Sept. 2026.

ExplainX. "Goal — Claude Code Slash Command Reference." ExplainX AI Documentation, May 2026, explainx.ai/slash-commands/claude-code-goal. Accessed 26 Sept. 2026.

MindStudio. "What Is the /goal Command in Claude Code? Autonomous Long-Running Tasks." MindStudio AI Blog, 12 May 2026, www.mindstudio.ai/blog/claude-code-goal-command-autonomous-tasks. Accessed 26 Sept. 2026.

Prospere AI. "Claude /goal Command: What It Does and How to Use It." Prospere AI Tech Insights, 21 May 2026, prospere.ai/en/blog/claude-goal-one-command-forces-ai/. Accessed 26 Sept. 2026.