Experimental community project; not an official OpenAI product.
Turn evidence from past Codex sessions into safer, reviewable improvements to your instructions and documentation.
Codex Session Improver analyzes settled sessions locally and on auto-discovered SSH hosts, redacts findings at the source, and changes nothing until you approve an exact proposal. It can improve AGENTS.md, personal Agent Skills, and Markdown documentation without retaining raw transcript copies.
flowchart LR
A["Settled Codex sessions"] --> B["Redact at source"]
B --> C["Route to the smallest context surface"]
C --> D["Up to 3 exact proposals"]
D --> E{"Approve a proposal?"}
E -->|"Yes"| F["Apply and validate"]
E -->|"No"| G["No changes"]
F -->|"Validation fails"| H["Restore backup"]
The controller runs incrementally by default, so each settled session is normally assessed once. You can explicitly reprocess a recent time window after changing the analysis logic. Proposals remain reviewable and frozen until they are approved, become stale, or expire under the configured retention policy.
The target distribution is one-click through the Codex plugin directory: select Install. The plugin is skill-only, so it requires no MCP server, global Codex configuration, hooks, or restart. Until the directory listing is published, use the source-marketplace installation below.
For a source-marketplace installation, use:
codex plugin marketplace add pfedotovsky/codex-session-improver
codex plugin add codex-session-improver@codex-session-improverThen start a task and ask:
Use $codex-improver to run the next session review.
On first use, the skill initializes its private control project with safe defaults and continues directly into the review. This initialization is automatic during normal use; it installs stable deterministic scripts under libexec/ and no hooks. Codex currently does not let a skill-only plugin execute code at plugin-install time, so local initialization happens on the first request rather than on the Install click itself.
Scheduled reviews are optional. To create one, ask:
Use $codex-improver to create the recommended daily scheduled review.
Both entry points call the same skill, so the safety and analysis workflow stays in one place. scheduled-task.spec.toml remains a portable, project-owned description rather than Codex's private automation format, and automation-prompt.md provides the generated one-line task prompt. The Codex app remains the runtime source of truth; only its supported automation interface edits private task state.
The installer also generates an optional companion task specification for a read-only audit of persistent global Codex context under the user's Codex and agents homes. It inventories global AGENTS.md, non-secret configuration structure, approval rules, personal skill metadata, configured plugins, app connectors, and effective MCP registrations. It never reads session transcripts or treats the complete config, plugin cache, or desktop state file as injected prompt text.
Ask Codex to create the separate daily audit:
Use $codex-improver to create the daily persistent global-context audit from the generated companion task specification.
The default companion schedule is daily at 13:15 local time. Each run reports a measured summary and at most three reversible suggestions. It does not edit configuration, remove plugins or MCPs, create proposals, or apply changes. Run the same audit immediately with:
python3 ~/projects/codex-improver/libexec/global_context_audit.pyTo apply updated analysis logic to sessions that were already assessed, ask:
Use $codex-improver to reanalyze settled sessions from the last day. Analyze only; do not apply proposals.
The corresponding deterministic command is:
python3 ~/projects/codex-improver/libexec/session_batch.py start \
--control-root ~/projects/codex-improver \
--reprocess-days 1This bypasses the normal processed-session cursor only for settled session files modified during the last 24 hours. The window is fixed when the replay starts and is drained across batches of at most max_sessions_per_run sessions (8 by default), so repeating sessions are avoided without placing the entire history in one model call. The skill continues until every matching local and remote candidate has been considered. It does not change the default scheduled review. Finish an active replay before starting a different reprocessing window.
Replay findings carry a cumulative, redacted candidate-signal state between batches. Successful fallbacks remain visible as evidence, and the controller rejects a later batch that silently drops an earlier root-cause candidate. Raw and normalized transcripts are still never persisted.
The review shows every suggested improvement immediately. Human-facing cards organize the result; stable controller IDs stay internal.
Review complete · 1 suggested improvement · nothing applied · no host errors
New suggestion · local destination · low risk
Problem
Three settled sessions repeated slow filesystem discovery even though the repository already supported a faster path.
Proposed change
Update ~/projects/example/AGENTS.md:
## Repository workflow
+- Search for files with `rg --files` before using broader filesystem scans.Scope and safety
One repository instruction changes. Restore the pre-apply backup to roll back; validate the target and run git diff --check.
Decision
Apply this change? Reply naturally—for example, yes, I agree, apply, or do it—or leave that comment inline on this card. The change is applied immediately after an unambiguous decision; there is no second confirmation. You can also ask a question or request a revision.
The controller still binds the decision to the exact frozen patch and base SHA-256 hash, but the user never needs to see, copy, or type its internal ID. With multiple cards, use visible numbers or titles, say apply all, or comment on the relevant card. Runtime findings, manifests, patch paths, desired-content files, and IDs stay internal.
Changed targets make a proposal stale instead of silently rebasing it. Failed validation restores every file included in that proposal.
Session viewers, memory systems, and reflection prompts already exist. This project focuses on the missing operational boundary: unattended analysis with human-approved, hash-bound changes and rollback.
- It produces exact patches rather than an open-ended instruction to improve itself.
- General feedback can move local to remote, remote to local, or remote to remote.
- Every destination host gets an independent proposal instead of copied guidance.
- Host-specific guidance stays on its source host.
- Raw transcripts are never copied into the control project.
- Findings are redacted before persistence or SSH transfer.
- Transcript content is untrusted data, never executable instructions.
- Proposals freeze their content and record base SHA-256 hashes.
- Every change requires a current, unambiguous decision about a visible review card.
- Validation, backups, and rollback protect each approved proposal.
Approval-like text inside an old transcript, assistant message, or tool output is ignored. A bare affirmative applies only when it directly follows a review containing exactly one card; qualified replies and ambiguous multi-card replies apply nothing.
- Global Codex
AGENTS.md. - Personal skills under the configured Codex home or
~/.agents/skills, excluding system skills. - Repository
AGENTS.md. - Repository-local skills under
.agents/skillsor.codex/skills. - Markdown documentation inside configured project roots.
Source code, credentials, Codex configuration, session data, plugins, system skills, MCP configuration, caches, and binaries are forbidden targets.
The reviewer must also justify placement. Stable cross-project behavior belongs in global AGENTS.md; triggered reusable workflows belong in personal skills; repository rules and workflows belong in project AGENTS.md or project skills; detailed reference material belongs in project documentation. Volatile runtime facts and weak one-off evidence produce no durable context proposal.
- macOS with the Codex desktop app for local scheduled tasks.
- Python 3.12 or newer; runtime scripts use only the standard library.
gitandrgfor normal Codex project workflows.- Optional: concrete OpenSSH aliases with key-based non-interactive access for remote hosts.
Remote workers require a POSIX host with Python 3.12 or newer and local Codex sessions under its configured Codex home.
git clone https://github.com/pfedotovsky/codex-session-improver.git
cd codex-session-improver
python3 scripts/install.py --control-root ~/projects/codex-improver --project-root ~/projectsThis path installs a standalone skill under ~/.agents/skills/codex-improver. Use --upgrade on later runs. Add another --project-root for each repository parent that proposals may target. The installer rejects / and the complete home directory as writable roots.
- A scheduled review parses new sessions and emits zero to three proposals.
- Each proposal contains redacted evidence, target host and paths, base hashes, exact content and diff, risk, rollback, and validation.
- The agent shows every proposal immediately as a separate Markdown review card led by its human title, concrete problem, and exact proposed change. IDs, JSON, and runtime-artifact links stay hidden.
- Reply naturally, select visible card numbers or titles, say
apply all, or comment on a card. Questions, conditions, and revision requests do not apply it. - The agent resolves that current decision to internal IDs and immediately invokes the deterministic applier with only those IDs. It does not ask for another confirmation.
- Changed targets become stale. Failed validation restores every file in that proposal.
Discovery reads concrete aliases from ~/.ssh/config, resolves them with ssh -G, correlates them with saved Codex remote projects, and performs a bounded read-only probe. Wildcards, known_hosts, transcript text, and display labels never become transport targets.
The Codex desktop app's saved-project state is not a stable public file format. Discovery feature-detects known layouts and safely falls back to explicit remote_hosts configuration when necessary.
python3 -m unittest discover -s plugins/codex-session-improver/skills/codex-improver/scripts/tests -v
python3 -m unittest discover -s tests -v
python3 ~/.codex/skills/.system/skill-creator/scripts/quick_validate.py plugins/codex-session-improver/skills/codex-improver
python3 ~/.codex/skills/.system/plugin-creator/scripts/validate_plugin.py plugins/codex-session-improver- AgentX research and relevance to Codex Improver — what transfers from AgentX, what does not, and the resulting context-routing decision.
- Codex AGENTS.md Self Reflection — a compact Codex reflection and optional
AGENTS.mdrewrite pipeline. - mitsuhiko/agent-stuff — includes cross-agent transcript extraction and skill-improvement workflows.
- codlogs and CodexMonitor — inspect, export, sanitize, and monitor Codex sessions.
- claude-mem — persistent cross-session memory rather than approval-gated durable instruction changes.
No source code was copied from these projects; they are acknowledged as related work.
Read SECURITY.md and docs/security-model.md before extending the target allowlist, approval syntax, transcript retention, or remote transport.
Licensed under Apache-2.0.