Skip to content

Latest commit

 

History

89 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Production Agentic AI

A research-oriented engineering handbook for designing and validating reliable production AI systems.

What this is

This is a focused research and reference repository on production agentic AI reliability. It centers on two primary themes, with failure modes treated as a lens across both:

  • Agent architecture, including authority and intent as research hypotheses
  • Context engineering and lifecycle management

Response delivery is a production failure case used to examine reliability trade-offs, not a third standalone theme. The repository contains explicit evidence boundaries and two narrow, tested local reference slices: the adaptive-response-filter protocol example and the bounded authority evaluator.

What this is not

  • Not a production platform, framework, or agent runtime.
  • Not a live CRM, EMR, governance, or multi-agent system.
  • Not a collection of enterprise design sketches; those are retained under archive/.
  • Does not claim production measurements or deployed results for the research notes.

Current Implementation Boundary

The implemented slices include a local authenticated-envelope/reassembly protocol example, a bounded authority evaluator with allow/deny/refer outcomes, an incremental NDJSON decoder, a browser harness for an NDJSON-over-HTTP candidate, and a deterministic CRM approval workflow simulation. The CRM workflow uses in-memory mock records and caller-supplied identities; it is not a live CRM, identity provider, model integration, durable approval service, or production runtime.

The active runtime and dependency boundary is documented once in ROADMAP.md and 04-reference-implementation/README.md.

Research Trail and Evidence Levels

Each investigation should connect the engineering question to its architecture, decision, executable slice, tests or benchmark, evidence record, known limits, and next validation. These levels describe the strongest evidence currently present for a specific slice; they do not rate production readiness.

Level Meaning
L0 — Conceptual Analysis, hypothesis, or proposed architecture; no executable behavior for the claim
L1 — Executable A local reference implementation exists
L2 — Reproducible Local tests or benchmarks reproduce the bounded behavior under stated conditions
L3 — Environment validated Tested against representative external systems in a bounded non-production environment
L4 — Operational Deployed with telemetry, failure handling, and named operational ownership
L5 — Production validated Evidence from an actual production or customer setting, with scope and conditions recorded
Investigation slice Current level Evidence trail and boundary
Response delivery L2 Engineering question and implementation boundary → ADR-001 → reference implementation and tests → browser evidence and limitations. Measurements are synthetic and runtime-dependent; representative CRM payloads, a second browser engine, and client decompression costs remain open.
Context lifecycle L0 Lifecycle question and hypothesis. No context-lifecycle implementation or benchmark is currently present.
Authority policy evaluator L2 Authority and intent question → bounded evaluator → tests. This local evaluator does not establish authenticated identity or semantic alignment.
CRM approval workflow L2 Proposed architecture → approval decision → in-memory simulation and tests → evidence and limitations. The simulation is serial and uses caller-supplied identities and in-memory state; concurrency, sink-failure recovery, and external integration remain unvalidated.

No current investigation slice demonstrates L3 or higher. The highest-value next evidence is a bounded non-production validation using real identity, conditional CRM writes, durable audit, rollback, telemetry, and named operational ownership; it is a proposed next step, not an existing capability.

Repository Capability Statement

Category Status in this repo Evidence
Executable implementation Local Python reference slice for adaptive response delivery Tests, demo, and code under 04-reference-implementation/adaptive-response-filter
CRM approval workflow simulation In-memory proposal, human-review, authority recheck, mock update, and JSONL audit sequence approval_workflow.py and focused tests; no live CRM, authenticated identity, or durable state
Research articles Documented engineering analysis and trade-off discussions Markdown articles and diagrams in this repository
Research hypotheses Context lifecycle and authority/intent analyses No context-lifecycle benchmark; a small time-bounded authority policy evaluator exists but does not implement semantic alignment
External / absent EMR, Spark, PostgreSQL/RDS pipelines, production deployments, and cloud service stacks referenced in archived examples Not present in the checked-in repository; not measured here

Featured Investigations

  1. Adaptive Response Delivery — Investigation, reference implementation and tests, and browser evidence. The strongest empirical slice here, with synthetic, runtime-dependent results and clear open validation questions.
  2. Context Lifecycle — Research note on what information should survive across interactions. Currently L0: no local implementation or benchmark.
  3. Agent Authority — Authority and intent research, alongside a bounded policy evaluator and tests. The evaluator is local and deterministic; it does not establish authenticated identity or semantic alignment.
  4. CRM Operational Copilot — Proposed architecture and approval workflow simulation, with its evidence boundaries. Architecture remains proposed; the workflow is an in-memory local simulation, not a live CRM integration.

Supporting artifacts: NDJSON decoder, referral tests, and CRM workflow tests.

The out-of-scope EMR-to-PostgreSQL architecture analysis is retained in the design archive.

Historical design sketches, evaluations, and dated review records are retained under archive/ and are not current project guidance.

How to Review This in 30 Minutes

This is a proposed design plus local simulation, not a deployed CRM product. Use this path to inspect the design and verify one safety control:

  1. Read the solution overview for the problem, scope, intended outcome, and what is not implemented.
  2. Open the C4 container diagram and follow the context and containers notes. The identity provider, CRM, model provider, approval queue, authority gateway, and audit store are proposed components.
  3. Read the ADR index and the authority options analysis to see alternatives, trade-offs, and provisional choices.
  4. Scan the STRIDE threat model, especially the flow IDs for model prompt injection, approval replay, and the recheck-before-write boundary.
  5. Run python prototypes/crm_operational_copilot/approval_workflow.py and python -m pytest -q prototypes/crm_operational_copilot/test_approval_workflow.py tests/test_authority_referral.py. The simulation uses in-memory mock records; 16 workflow test cases cover approval, binding, current requester/reviewer authority, stale records, expiry, and replay, while 7 referral tests cover evaluator, aggregate, session, and ticket behavior.

For deeper discovery, see the NFR and sizing worksheet, AWS deployment candidate, cost model, and customer questionnaire. Their assumptions and numbers are provisional, not requirements or quotations.

Repository structure

Quick Start

# Install development dependencies
python -m pip install -e ".[dev]"

# Run tests (pytest testpaths are configured in pyproject.toml)
python -m pytest -q

# Run lint
python -m ruff check .

# Run type checks
python -m mypy 04-reference-implementation/adaptive-response-filter 04-reference-implementation/authority_policy.py tests/test_authority_policy.py

# Run the authority policy evaluator tests
python -m pytest -q tests/test_authority_policy.py

# Run the local approval workflow simulation and focused tests
python prototypes/crm_operational_copilot/approval_workflow.py
python -m pytest -q prototypes/crm_operational_copilot/test_approval_workflow.py

# Run referral outcome tests
python -m pytest -q tests/test_authority_referral.py

# Run the NDJSON stream decoder tests
python -m pytest -q tests/test_ndjson_stream.py tests/test_ndjson_guards.py

# Run the reference demo
python 04-reference-implementation/adaptive-response-filter/demo.py

# Run the local evidence benchmarks
python benchmarks/response-delivery/benchmark.py

# Run the local browser delivery experiment
python benchmarks/response-delivery/browser_benchmark.py

Evidence posture and benchmark layer

This repository is intentionally a reliability-first, architecture-first knowledge base and local reference implementation. It is not a framework, production runtime, or deployed AI platform.

The evidence layer under benchmarks/README.md contains bounded local measurements and policy tests, not production deployment evidence. The response-delivery browser harness compares full JSON, gzip full JSON, NDJSON, and gzip NDJSON over server-paced loopback. It has three distinct artifacts: a legacy three-mode standalone HeadlessChrome run, a clean four-mode standalone HeadlessChrome run with a passing Long Task positive control and recorded harness provenance, and a four-mode Electron-embedded run whose Long Task control failed and whose historical provenance is incomplete. The clean standalone and Electron four-mode runs use matching payload sizes; their runtime-specific results should not be conflated. Detailed timing and compression-level results are in the browser matrix evidence.

The repository's strongest local evidence is currently:

These artifacts are intentionally narrow and reproducible. They are not presented as production telemetry, production incident data, or deployed-system benchmarks.

The repository's evidence model is:

  • implemented and tested locally: real code + tests + CI
  • conceptual and architecture-level: reasoning and design intent
  • measured locally under controlled conditions: response-delivery implementation overhead and the paced standalone-browser harness
  • no context-lifecycle behavior is currently implemented or tested
  • future work / not yet evidenced: production deployment, deployed browser client, or fleet-scale data

Repository Packaging and Dependency Reality

This repository is intentionally not presented as a reusable installed Python package API. The editable-install metadata in pyproject.toml declares a lightweight project config without claiming a production runtime surface.

This matters for reading the project correctly:

  • The actual executed code is the repository's local source tree, especially the modules under 04-reference-implementation/adaptive-response-filter
  • The repository does not claim to ship a production Python package, a web framework, or a runtime service.
  • The project is a documentation and reference-workspace repository first, with a small local implementation slice second.

How to Read this Repository

This repository is a documentation-led engineering knowledge base. Its primary product is the article narrative, the architecture reasoning, and the evidence trail behind specific engineering decisions.

Read the articles as engineering arguments and research records, not as API documentation for a complete product. Each article describes a production problem, compares possible approaches, and records the trade-offs and evidence behind a decision.

Use the diagrams as architecture-level explanations of the system being discussed. A diagram may describe a production deployment, a proposed design, or a local simulation that is not present in the code.

The reference implementation is an isolated, pedagogical Python slice. It makes selected policies, authenticated envelopes, session recovery, and reassembly contracts executable, but it is not an operational Agentic AI service: it has no model runtime, database, browser client, or serving stack.

Important boundary rule: the repository is not a live AI product, not a deployment environment, and not a monorepo for a production stack. It is a bounded, inspectable reference implementation and a focused research archive.

Architecture & Presales Evaluation

Any technical or customer-facing evaluation of this repository must follow the evidence-first methodology in prompts/architecture-presales-evaluation.md.

That document requires clear separation of:

  • what is actually present and executable in the repository
  • what is only documented
  • what is design intent or hypothesis
  • what cannot be verified from the reviewed repository

It also prevents unsupported customer claims. Research articles and architecture sketches remain research records; they are not automatically production capabilities.

A useful reading sequence is:

  1. Start with the article's problem, hypothesis, evidence, and decision to understand the engineering question.
  2. Read its diagrams as the article's conceptual architecture, checking the surrounding text for what is measured, what is illustrative, and what is proposed but not implemented.
  3. Open the linked reference modules to inspect only the executable subset; do not infer that an article's production workflow is fully represented by the code.
  4. Run the local tests and demo to verify the behavior of that subset. The tests do not reproduce the article's production benchmarks unless the article explicitly says they do.
  5. Consult ROADMAP.md for the repository scope and the distinction between current dependencies, implemented packages, and future research context.

The repository should be read in three layers:

  • Layer 1: articles and engineering analysis
  • Layer 2: diagrams and architecture narratives
  • Layer 3: local Python reference implementations and their tests

The layers complement each other, but they are not interchangeable. A production article may discuss a gateway, streaming UI, or data plane that is not present in the local code; the repository is intentionally bounded.

Systemic Context vs. Active Dependencies

The repository has two different kinds of material and they should not be conflated.

Active dependencies in the local execution boundary:

Systemic context or future-facing ideas that may appear in the article narrative:

  • FastAPI or other web-service runtimes
  • Pydantic validation layers
  • Redis, Postgres, or other persistence queues
  • browser client logic and TypeScript transport adapters
  • Managed LLM or observability SDKs
  • deployment and orchestration infrastructure

These broader possibilities are valid research ideas and architecture vocabulary, but they are not required to run the repository's local tests, demo, or lint checks. The current execution environment is intentionally small.

The current local reference slice is intentionally narrow. It can be exercised by the checked-in tests and demo without requiring external services. Its retry and fallback callbacks model local session behavior; they are not connected to a real transport.

License

See LICENSE.

About

Building Reliable Agentic AI Systems

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages