Skip to content
View flam7791's full-sized avatar
  • Paris, France

Block or report flam7791

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
flam7791/README.md

Flavio

I work on enterprise AI adoption and governance: moving AI from pilots to governed, measurable services in large organisations. These repositories show how I approach it in practice: a framework that makes AI services consistent, and the working systems it is built from. Small, tested, auditable, explicit about what they measure and what they cost, and able to run entirely on open-weight models on an organisation's own infrastructure.

The framework

ai-engineering-framework: how an organisation identifies, builds, industrialises and operates AI solutions the same way every time. A reference architecture with on-premises, hybrid and cloud topologies; six solution patterns, each with a working reference implementation below; engineering standards that run as checks in CI (aief check); a use-case intake that scores proposals on feasibility, information sensitivity, security, cost, scalability, interoperability and sustainability and recommends a pattern and topology (aief intake); lifecycle gates and a handover pack; and a service template that starts every new project local-first, evaluated and production-ready.

Reference implementations, by stage

Stage Repository What it shows
Identify governed-agents (use-case triage desk) Agents that register a proposed AI use case, assess risk and cost, choose a pattern and submit a decision record for sign-off
Build policy-evidence-mcp An MCP server giving assistants cited access to official statistics (SDMX) and policy documents: hybrid RAG with a sensitivity ceiling, per-caller access from bearer tokens or Entra ID app roles, SharePoint libraries synced through Microsoft Graph with Purview labels as the classification, retrieval evaluation as a CI gate
reference-resolver-agent Deterministic scoring first; the model chooses only among records actually retrieved; a bounded search agent; uncertain cases to a human review queue; measured on a gold set
oecd-data-pipeline Python computes every figure, a model (Microsoft 365 Copilot or a local open-weight model) writes only the wording, and a validator rejects any note with a number not in the data
copilot-team-knowledge A verified knowledge layer that Microsoft 365 Copilot answers from: only active cards at or below a classification ceiling are published, to SharePoint through Microsoft Graph; also answerable by a local model
Industrialise governed-llm-gateway One door to every model: routes each request to the cheapest adequate model (Claude, Azure OpenAI with Entra ID, or local), masks personal data, enforces budgets, chargeback without storing content, OpenTelemetry traces
governed-agents Multi-agent runtime where a policy engine decides every tool call: least privilege, autonomy levels, four-eyes approval for external actions, audit trail, kill switch, trajectory evaluation with a planted prompt injection, one OpenTelemetry trace per run
Operate governed-ai-platform The components as one operable service, on Docker Compose or Kubernetes (restricted pod security, default-deny network policies): Prometheus alerts, one trace per agent run across the services, pinned versions, runbook, an end-to-end test in CI, and a sovereign mode in which no external model exists

Local and open-weight by design

Every system runs without a commercial API: Ollama on a laptop, or any OpenAI-compatible server (vLLM on a GPU server, for example), with commercial models as a governed choice through the gateway. Measured with Llama 3.1 8B on a laptop CPU, answers recorded and replayed in CI: the resolver matched Claude (precision and recall 1.00) at zero cost; a service generated from the framework passed 10 of 10 cases, including one where the model followed a planted instruction and the validator withheld the answer; the knowledge layer passed 11 of 12 with no blocking failure. The first runs also exposed two integration bugs, now fixed and tested (details).

How I build

  • Deterministic where possible, models where they add value, people where it matters
  • Guardrails enforced in code, not only in prompts
  • Every system ships with an evaluation, a cost figure and a way to run it locally
  • Standards that run: each repository is checked against the framework's standards in CI

The design decisions behind each project are documented in its docs/ folder.

Python · MCP · RAG · AI agents · FastAPI · pandas · SDMX · Claude API · Azure OpenAI · Microsoft 365 Copilot · SharePoint · Ollama · open-weight models · Microsoft Graph · Entra ID · Docker Compose · Kubernetes · Prometheus · OpenTelemetry · GitHub Actions · Copier

Paris · English, Italian, Spanish, French (working knowledge)

Pinned Loading

  1. ai-engineering-framework ai-engineering-framework Public

    A local-first engineering framework for governed AI services: reference architecture, patterns, standards as CI checks, use-case intake and a service template.

    Python

  2. governed-agents governed-agents Public

    Governed multi-agent runtime: a policy engine decides every tool call, external actions wait for a named approver (four eyes), every step is audited, and agents are evaluated on what they did, not …

    Python

  3. governed-ai-platform governed-ai-platform Public

    Reference deployment of a governed AI stack: LLM gateway, MCP evidence server and multi-agent service as hardened containers, with secrets, Prometheus monitoring and alerts, pinned versions, a runb…

    Python

  4. governed-llm-gateway governed-llm-gateway Public

    Internal LLM gateway: routes each request to the cheapest adequate model (Claude, Azure OpenAI or a local open-weight model), masks personal data, enforces team budgets and reports chargeback witho…

    Python

  5. policy-evidence-mcp policy-evidence-mcp Public

    MCP server giving AI assistants cited access to official statistics (SDMX) and policy documents (hybrid RAG: keywords plus embeddings), with a sensitivity ceiling and a retrieval evaluation as a CI…

    Python

  6. reference-resolver-agent reference-resolver-agent Public

    Agentic workflow that resolves messy bibliographic references to DOIs: deterministic scoring first, a bounded Claude agent for the rest, a human review queue when in doubt. Measured on a gold set.

    Python