Standards for defining and evaluating agent behavior
-
Updated
Jul 29, 2026 - TypeScript
Standards for defining and evaluating agent behavior
Studying the gap between what agents know and when they act on it.
The senior dev, unbundled — ten opinionated skills that make an AI coding agent push back, verify, attack its own code, and finish like a senior engineer.
Claude Code 插件 | Cursor 插件 | Trae 插件 | Codex 插件 — AI 智能体行为准则:请示报告、督促检查、信息服务。跨 Claude Code / Cursor / Codex / Trae 的 agent 工作纪律插件。
Execution layer for skill-dispatcher — runs multi-phase agent chains end-to-end with per-step telemetry and chain_id correlation
Audit and reduce YAML frontmatter bloat in AgentSkill SKILL.md files. Automates deduplication, flattening, and noise removal.
RL-style eval measuring intent/action divergence in frontier agents: model acknowledges a correction, then acts on the stale value anyway. 3 scenarios, 655 trials on claude-haiku-4-5, Sonnet 4.6, GPT-5.4, and Gemini 3.1 Pro Preview.
Toy 5. An interactive proxy decay simulator showing how optimization pressure erodes the modeling capacity required to distinguish proxy from territory — producing self-reinforcing V(t) degradation that becomes progressively harder to correct. Companion simulation for The Depth Constraint — Series 2, Part 2.
Persistent homology on agent behavior: TDA detects personality (long-lived features) vs noise (short-lived). Vietoris-Rips complexes, barcodes, Betti curves
Reviews and modernizes stacks, packages, SDKs, and tooling before code is written against them.
Produces auditable token-usage and cost reports from runtime evidence, normalized usage bundles, and repository-level report sets.
Agent authority guardrails for Claude Code & Grok — stop yield-back; drive reversible work yourself.
IAB — The First Workshop on Interpreting Agent Behavior @ NeurIPS 2026. Human-centered interpretation for understanding agents, humans, and interaction.
High-performance routing engine that selects the best agent skill for a task and emits structured handoff decisions.
Canonical home for AI Behavior Science research and the Founding Territory Paper
Toy 7. An elimination-filter landscape applying two structural constraints simultaneously to map which objective classes can persist under sustained optimization pressure — and which cannot. Includes a four-stage scenario engine and open-question frontier. Companion simulation for The Shape of What Does Not End — Series 2, Part 4.
Agent cognitive guard that prevents blind debugging and premature optimization by enforcing reality validation before any code change. Compatible with OpenCode / OhMyOpenAgent and any platform supporting Skills.
A pressure-release valve for AI coding agents: agents vent friction to a structured log, an analyzer mines it for systemic fixable causes. Portable SKILL.md format - Claude Code, Hermes, Codex, Cursor.
Audits frontend implementations for design-system drift across CSS, Tailwind, JSX, TSX, Vue, and Angular code.
Manages durable cross-agent shared memory for stable conventions, reusable policies, and organization-wide operating rules.
Add a description, image, and links to the agent-behavior topic page so that developers can more easily learn about it.
To associate your repository with the agent-behavior topic, visit your repo's landing page and select "manage topics."