Codex Security CLI Tutorial: Scan Your Repository for AI-Found Vulnerabilities
Learn how to install and use the Codex Security CLI for repository scans, Git diffs, deep analysis, cost limits, pre-commit checks, and CI workflows.
On this page
The current Codex Security CLI can scan a local repository, validate potential vulnerabilities, preserve evidence and coverage data, and prepare fixes for review without requiring you to wire a separate security scanner into your project. The public package now supports standard and deep scans, targeted paths, Git diffs, working-tree checks, structured output, cost limits, and a pre-commit hook, making it useful as an AI-assisted security step rather than just a one-off experiment.
This tutorial walks through the local CLI workflow from installation to the first scan, then shows how to narrow the scope, inspect the evidence, add repository-specific security context, and use the scanner before commits or in continuous integration. The commands are based on the current Codex Security CLI documentation and package instructions.
What Codex Security actually does during a scan
Codex Security is designed to investigate potential vulnerabilities rather than simply match files against a fixed list of signatures. It builds a threat model for the project, examines code paths and attacker entry points, investigates suspicious behavior, and attempts to validate potential findings. When a finding is validated, the system can propose a patch for a developer to review rather than silently changing the repository.
That distinction matters because an AI security scan is only useful when you can inspect why something was reported. Current scan output includes finding details, evidence, remediation information and coverage data. Coverage can also identify deferred areas or open questions, so an apparently clean result should not automatically be interpreted as proof that every part of a large codebase was exhaustively reviewed.
Install the Codex Security CLI with the current runtime requirements
The public package is distributed as @openai/codex-security. The current quickstart requires Node.js 22.13.0 or later, with Node.js 24 and 26 also documented as supported, while Python 3.10 or later is required for scans and several other operations. You can use npx without maintaining a separate global CLI installation.
npm install @openai/codex-security npx @openai/codex-security --version npx @openai/codex-security info --jsonThe version command confirms that the executable can start, while info --json gives you package and bundled-plugin information. Checking these before a long scan is worthwhile because security tooling depends on several moving parts, and a reproducible environment makes later scan comparisons much easier.
Authenticate before you scan code you are authorized to assess
For local interactive use, the CLI supports ChatGPT authentication. Automated environments can use an OpenAI API key instead. Remote machines can use device authentication, which avoids depending on an interactive browser session on the machine running the scan.
npx @openai/codex-security loginFor an automated environment, provide the credential through the environment rather than putting it in a command that could be recorded in shell history.
export OPENAI_API_KEY="your-api-key-here"The scanner operates with the permissions of the account running it, so authentication does not create a separate security boundary. Only scan repositories you own or have explicit permission to assess, and remove unrelated credentials from the environment before starting a local scan.
Do not point Codex Security at an untrusted repository just because you want to see what it finds. Local scans can operate with the permissions available to the account running the tool, so the repository, scanner configuration and surrounding environment need to be treated as trusted inputs.
Run a dry run before spending tokens on a real scan
A dry run is the safest first test because it checks local inputs without starting the actual Codex analysis. It can catch problems with the repository, selected paths, knowledge-base files and output location before you spend time or model budget on the scan.
REPOSITORY=/path/to/repository SCAN_DIR=/path/outside/repository/codex-security-results
npx @openai/codex-security scan "$REPOSITORY"
--output-dir "$SCAN_DIR"
--dry-runKeep the output directory outside the repository and outside its enclosing Git worktree. Scan results can contain source excerpts and vulnerability details, so storing them in a deliberately private directory also makes retention and access control easier to manage.
Run the first standard Codex Security scan
Once the dry run passes, remove --dry-run and run the standard scan. Standard mode is the appropriate starting point because it gives you the normal repository review without immediately expanding the investigation into a deeper scan.
npx @openai/codex-security scan "$REPOSITORY"
--output-dir "$SCAN_DIR"The command produces a completion summary and saves the detailed results in the selected directory. The current workflow can produce a readable report.md together with structured files such as findings.json and coverage.json. The useful question after the command finishes is not simply how many findings appeared, but whether each finding has enough evidence to reproduce and evaluate the alleged problem.
Read the report and coverage separately
Open report.md first because it is the easiest way to understand the scan as a human reviewer. Then inspect the structured finding and coverage files when you need to automate the results or compare scans over time. A finding can contain severity, confidence, locations, evidence and remediation information, while the coverage record explains which surfaces were reviewed and whether work was deferred.
This separation prevents a common mistake with AI security tools: treating the final finding count as the entire result. A scan with zero findings but incomplete coverage is different from a scan with zero findings and complete coverage. Likewise, a high-severity finding with weak evidence should be investigated differently from one that the validation stage reproduced convincingly.
Use targeted scans when the whole repository is unnecessary
You do not have to scan every directory on every run. The CLI can restrict a scan to selected paths, which is useful for a monorepo where different services have different owners or security boundaries.
npx @openai/codex-security scan "$REPOSITORY"
--path services/billing
--path packages/authThis approach makes the result easier to reason about because the scanner is examining explicitly selected areas. It also gives teams a practical way to run focused checks when a developer is changing one sensitive component rather than waiting for a repository-wide analysis.
Scan Git changes when you care about what just changed
For development workflows, scanning the entire repository can be unnecessary. Codex Security supports committed-change scans and working-tree scans, allowing the review to concentrate on changes instead of repeatedly analyzing unrelated code.
To compare committed changes against a base revision, use a Git diff target:
npx @openai/codex-security scan "$REPOSITORY"
--diff origin/main
--head HEADFor staged and unstaged changes, use the working-tree mode:
npx @openai/codex-security scan "$REPOSITORY"
--working-tree
--base HEADThese modes are mutually exclusive with path selection, and the repository argument needs to point to the Git worktree root. The selected revisions also need to exist locally. That makes a diff scan particularly useful before a pull request because the result is tied to the change rather than the entire history of the application.
Add SECURITY.md when the scanner needs your application's threat model
Generic security rules cannot fully describe an application. A payment endpoint, an internal administration panel and a public image-processing service can have very different trust boundaries and acceptable behaviors. Codex Security supports repository security guidance through SECURITY.md, including information about the threat model, security invariants, reportable findings and exclusions.
You can also ask the CLI to draft security-policy guidance for a repository or component:
npx @openai/codex-security policy.The generated policy is a draft for review; it is not something you should blindly install. The useful part is the opportunity to make assumptions explicit. If your application treats one database table as highly sensitive or considers a particular network boundary security-critical, recording that context gives future scans information that source code alone may not communicate clearly.
Use deep mode only when the extra investigation is justified
Codex Security has a deeper scan mode for repositories or paths that need broader investigation. It can use multiple workers and additional discovery passes, so it is better suited to a deliberate security assessment than to every save or commit.
npx @openai/codex-security scan "$REPOSITORY"
--mode deep
--workers 2
--subagents 0
--stop-after-no-new 3
--max-discovery-runs 10
--max-time-hours 1.5The limits are there for a reason. A deep scan can consume more time and model usage than a standard review, while the worker and discovery settings determine how much investigation is attempted. Start with standard mode, examine what it misses or leaves uncertain, and then use deep mode where the additional analysis has a clear security purpose.
Put a cost ceiling around automated scans
AI security analysis has a variable model cost, so an automated workflow should not assume every repository costs the same to inspect. The CLI supports a maximum estimated cost for an individual scan.
npx @openai/codex-security scan "$REPOSITORY"
--max-cost 5The value is an estimated U.S.-dollar limit rather than a guarantee that every in-flight request will stop at exactly that amount. Current documentation notes that requests already in progress can finish slightly above the limit. If a deep scan reaches the limit after aggregating completed worker results, the saved report can be marked with partial coverage.
Turn the scan into a pre-commit security check
Once the basic workflow is producing results you trust, Codex Security can install a Git pre-commit check. The purpose is not to replace your existing tests or static analysis, but to add an AI-assisted security review to the point where developers are about to record a change.
npx @openai/codex-security install-hookThe current CLI documentation says the hook scans staged and unstaged changes and blocks high-severity findings and scan errors. That makes it more suitable for catching security issues before they leave a developer's machine, while a separate continuous-integration workflow can provide the broader repository or pull-request review.
Use structured output when the scan becomes part of CI
Human-readable reports are useful for developers, but automation needs predictable data. Codex Security can emit JSON and can evaluate a severity threshold for continuous integration workflows.
npx @openai/codex-security scan "$REPOSITORY"
--diff origin/main
--output-dir "$SCAN_DIR"
--json
--fail-on-severity high
> "$SCAN_DIR/codex-security.json"This changes the role of the scanner from an advisory tool into one component of a policy pipeline. The important distinction is that the severity threshold is your workflow's decision, not proof that the AI's classification is infallible. Keep the evidence and validation details available so a developer can investigate a failed build rather than receiving only a red status.
Review AI findings before accepting a patch
Codex Security can propose fixes and, in supported workflows, create a pull request for verified patches. That does not remove the need to review the change. A security fix can affect authentication, authorization, data handling or application behavior beyond the vulnerable line itself.
The safer workflow is to inspect the reported attack path, confirm the evidence, understand the proposed remediation, and then review the resulting code as you would any other security-sensitive change. If the scanner cannot establish the vulnerability convincingly, the right outcome may be further investigation rather than an automatic patch.
The practical workflow is scan, verify, then automate
Codex Security becomes more useful when its commands are assigned clear jobs. Use a dry run to validate the environment, a standard scan for routine coverage, targeted or diff scans for development changes, deep mode for investigations that justify additional analysis, and structured output when a CI system needs a machine-readable result. Keep the reports outside the repository and treat their contents as sensitive because they can include code excerpts and vulnerability details.
The final step should remain human review. An AI security scanner can investigate code and validate plausible attack paths, but the engineering team still decides whether the evidence is sufficient, whether a proposed fix preserves intended behavior, and whether a finding should change the release decision. That boundary is what turns the CLI from an interesting AI experiment into a usable part of a software security process.
Written by


