Independent researcher and builder working on trustworthy AI-agent execution.
I study what final-answer evaluation misses when agents use tools, cross authorization boundaries, change state, and create external effects. My work connects evaluation methodology with applied systems: the same concepts should survive both a controlled experiment and a production workflow.
- Trajectory-aware evaluation: comparing final-output judgments with independent process evidence.
- Delegated authority: representing approval, scope, identity, queueing, and permission to act as separate states.
- Failure-preserving research: keeping aborted runs, invalidated holdouts, corrections, and limitations auditable.
- Applied autonomy: building owner-governed systems and documenting the operational responsibility they create.
Nirmata is my independent pilot research program for evaluating AI agents from execution trajectories—not only final outputs.
| Explore | What you will find |
|---|---|
| Start with the project | Research map, current status, scope, and limitations |
| Methodology | Factorized evaluators, process taxonomy, blinding, and metrics |
| Evidence ledger | Supported statements beside unsupported extrapolations |
| Experiment registry | Exploratory, development, integration, and invalidated work |
| Articles | Accessible essays in English and Brazilian Portuguese |
Current evidence boundary: the evaluation pipeline has passed development and integration checks. No blind confirmatory result is claimed; the first planned holdout was invalidated before scoring after ground-truth exposure.
- Why final answers are not enough to evaluate AI agents
- Approval is not authorization in agentic systems
- A burned holdout is still research evidence
I prefer inspectable claims over impressive claims, explicit authorization over conversational implication, and preserved failures over cleaned-up histories.
Research materials are published in English and Brazilian Portuguese.
ORCID 0009-0007-1104-9204 · LinkedIn · Nirmata
Pesquisadora independente e construtora de sistemas voltada à execução confiável de agentes de IA.
Investigo o que a avaliação da resposta final não percebe quando agentes usam ferramentas, atravessam limites de autorização, alteram estados e produzem efeitos externos. Meu trabalho conecta metodologia de avaliação a sistemas aplicados: os mesmos conceitos precisam resistir tanto em um experimento controlado quanto em um fluxo de produção.
O Nirmata é meu programa piloto independente de pesquisa sobre avaliação por trajetórias de execução, autoridade delegada e metodologias que preservam falhas.
