Skip to content

How to design a CLI for AI agents

By Vincent Schmalbach · Updated

When an AI agent uses a command-line interface (CLI), it runs commands and reads their output to decide what to do next. A CLI designed only for people at a terminal can leave an agent guessing: it might print status messages where the agent expects data, wait for confirmation no one can provide, or return an error the agent cannot tell apart from another failure.

Design the CLI as a stable contract between the program and its callers. Give commands predictable names and return structured results. Document failure codes, and let callers preview and control changes. That stable contract also helps people who use scripts or need to automate a task.

Give commands a predictable shape

Use a consistent command pattern so an agent can discover actions and form valid commands from the help text. A noun followed by a verb is one workable convention: widget list, widget create, and widget delete. Keep flags consistent across commands, use descriptive long names such as --output or --dry-run, and avoid reusing a flag for different meanings.

The Elastic CLI design discussion emphasizes consistent flags, authentication, output, and failure behavior. Consistency reduces the number of exceptions an agent must handle. Document required arguments, accepted values, and defaults in each command’s help output. For larger CLIs, provide a way to inspect available commands and their parameters without invoking them. The agent CLI guide also recommends schema-based discovery, such as exposing command parameters in a machine-readable form.

Treat changes to command names, flags, and output fields as changes that can break existing callers. If an agent expects a JSON field named id, renaming it can break the caller even when the command still runs. Add fields without removing or renaming existing ones when you need to evolve the output format.

Separate data from diagnostics

Send machine-readable results to standard output, or stdout, and send warnings and error details to standard error, or stderr. This separation lets an agent parse a result without accidentally reading a progress message as data. Offer a consistent option such as --json rather than making the agent extract data from tables meant for people. The Elastic CLI guidance also discusses JSON Schema, a formal description of JSON fields and types, as a way to make output easier to inspect and validate.

Return documented exit codes so callers can distinguish outcomes. For example, a command can use separate codes for success, invalid input, missing resources, denied access, and conflicts. Keep the code meanings stable and explain them in help or command documentation. An agent can then respond to a conflict differently from a malformed command instead of treating both as generic failure.

Detect whether the CLI is connected to an interactive terminal, commonly called a TTY. In a noninteractive context, avoid prompts, pagers, and color-dependent output. A prompt can leave an agent waiting indefinitely, while terminal styling can interfere with parsing. If a command requires human input, provide an explicit option for supplying that input in automation, or fail with a clear error.

Make changes previewable and controlled

Give commands that create, update, or delete data a --dry-run mode. A dry run reports the intended action without making the change. Return a clear status in structured output and document its exit code, so the agent can distinguish a successful preview from a completed mutation. For destructive actions, require an explicit confirmation in interactive use and provide a documented, noninteractive way to authorize the action. Do not let unattended commands silently wait for confirmation.

Validate input before acting. Reject unknown fields, invalid values, and malformed structured input with an error that identifies the problem. When a command accepts JSON, publish its accepted shape or schema. Restrict agent access through permissions or command allowlists when the surrounding environment supports them. Store credentials through an operating-system keychain or environment-based configuration rather than asking users to place secrets directly in command arguments, where they can appear in shell history or logs. The Snyk guidance on securing agent workflows describes risks such as prompt injection and credential exposure; treat an agent as an untrusted caller and limit the resources its credentials can reach.

A dry run is useful, but it does not replace access controls or isolation. The agent still needs only the permissions required for its task, and the CLI must enforce those permissions when it executes a real change.

Keep results useful without flooding context

Agents have limited context for reading command output and deciding what to do next. Return the fields needed for the next step, and provide ways to narrow larger results, such as filters, field selection, or output templates. For example, a list command should let a caller request a small set of records or fields instead of printing every available detail. Keep error messages actionable: name the invalid argument or missing resource and, where possible, show the relevant command or accepted value.

For commands that take a long time, report progress or completion without mixing it into the machine-readable result. Keep human-facing progress on stderr when stdout carries JSON. Document whether the command returns only when it finishes or sends status updates as it runs. Callers need to know whether to wait, read a final result, or handle partial output. A CLI does not suit every integration. A software development kit (SDK), which lets code call a service directly, may suit callers that need more control over sessions or events. A CLI remains useful for bounded actions that callers can express as commands. The trade-off depends on the integration’s needs, as discussed in this comparison of CLI calls and SDK integration.

Test the contract, not just the command

Test each command in both interactive and noninteractive contexts. Check that JSON output parses and diagnostics stay on stderr. Verify that exit codes match the documented outcomes. Test that invalid input causes no mutation and dry runs leave state unchanged. Include examples in help that a person or agent can copy and run.

For Go CLIs, Rungrad provides a framework built on Cobra, a language-neutral behavior specification, and a scorer that tests a configured executable against relevant checks. The framework works with Go. CLIs written in other languages can use the specification and scorer separately. Rungrad does not provide an agent runtime or an MCP server, and its scorer is not a security certification. The scorer also distinguishes applicable checks from cases that were not configured. A result marked non-applicable does not count as a pass.

A generated starter offers a concrete way to inspect the contract:

rungrad new hello
cd hello
go mod tidy
go test ./...
go build -o hello .
./hello widget list --json
./hello widget delete alpha --dry-run
rungrad score ./hello --read "widget list" --mutate "widget create demo" --destructive "widget delete alpha" --update

In the starter, the dry-run command prints DRY RUN: would DELETE /widgets/alpha and reports that no changes were made. The preview uses predefined widget data stored in memory. Its text resembles an HTTP request, but the command does not contact an external API. Existing Cobra CLIs can adopt Rungrad features incrementally, as described in its migration guidance.