Skip to content

Require explicit premium model eligibility and surface turn cost - #348

Merged
witbrock merged 1 commit into
mainfrom
codex/JVNAUTOSCI-2624-premium-model-controls
Aug 5, 2026
Merged

witbrock merged 1 commit into
mainfrom
codex/JVNAUTOSCI-2624-premium-model-controls

Conversation

@witbrock

@witbrock witbrock commented Aug 5, 2026

Copy link
Copy Markdown
Member

Outcome

Implements JVNAUTOSCI-2624 so external generative models can run only when the exact provider/model is in the existing actor- and organisation-scoped allowed-model pool. Persisted, content-free provider usage and effective-model evidence now supports honest per-turn cost estimates in Thinking and bounded recent-cost summaries in Settings.

Smallest coherent design

  • Reuse enabled_llms as the single scoped execution authority; enforce it at the shared generation boundary and the few raw vision/live-probe call sites.
  • Keep requested, selected, and provider-reported effective model identities distinct.
  • Persist stable call IDs, provider-reported token usage, completeness, and pricing provenance; deduplicate retries and retain multi-call turns.
  • Estimate only from exact represented provider/model pricing. Unknown identity, unsupported usage, or incomplete evidence remains partial or unavailable rather than being guessed.
  • Show effective model, token use, call count, cost status, and lineage in Thinking; show a bounded actor-scoped recent summary in Settings.

Embeddings are intentionally outside chat-model eligibility: they are fixed non-generative infrastructure, while the RAG runtime LLM already passes through the gated generation boundary. Adding embeddings to enabled_llms would expand this task and break the existing retrieval contract.

Represented pricing evidence

  • Predicate #V#has_model_pricing_json: 40543e7b-e706-40db-8df4-b58608456a6e
  • Terra registry relation: 6a732a105c780564a2070522
  • Luna registry relation: 6a732a265c780564a2070523
  • Source: official OpenAI API pricing, version openai-standard-observed-2026-08-05

Validation

  • Backend affected-path suite: 231 passed
  • Adaptive-turn suite: 132 passed
  • Orchestrator and stale-response coverage: 53 passed
  • Frontend: 272 passed; static lint clean
  • Targeted cost, eligibility, history, identity-lineage, retry-deduplication, and multi-call checks passed
  • ruff, py_compile, and diff checks passed
  • Automated tests invoked no provider models

The existing repository-level Sol prohibition remains unchanged; this change neither invokes nor restores Sol.

Copilot AI lite review requested due to automatic review settings August 5, 2026 12:31
});

test('uses the persisted effective allow list, not staged edits', () => {
const eligibility = document.getElementById('openaiModelEligibilityStatus');

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Human review recommended

It changes core model-execution authority enforcement and cost/usage telemetry across multiple backend and frontend paths, warranting final human review despite strong test coverage.

Pull request overview

Implements scoped, explicit eligibility checks for running external/premium generative models (reusing enabled_llms as the execution authority) and adds content-free token/cost evidence so the UI can surface per-turn estimated costs (Thinking) plus bounded recent historical cost summaries (Settings).

Changes:

  • Add a shared backend eligibility gate (assert_model_execution_allowed) and wire it into structured tool calls, legacy provider calls, settings probes, and vision/image paths.
  • Persist and normalise richer per-call model identity lineage (requested/selected/effective + source), token usage (including cache breakdown), and compute a per-turn llm_usage_cost_summary.
  • Add an actor-scoped historical cost summary service + Settings endpoint and render eligibility/cost status in the Settings UI and Thinking card.
File summaries
File Description
tests/frontend/settingsOpenAiModelChange.test.js Updates Settings page expectations for eligibility/cost UI and “allowed” wording.
tests/frontend/settingsModelPoolControls.test.js Adds tests for persisted allow-list precedence and OpenAI cost-summary fetching/rendering.
tests/backend/test_von_turn_execution_debug_info.py Validates finalised debug payload includes cost summary and preserves paid calls on failure.
tests/backend/test_von_generate_workflow_instances.py Ensures model registry snapshot is passed through generate/adaptive turn and eligibility denials are recorded.
tests/backend/test_thinking_llm_call_timestamps_and_precedence.py Extends exchange normalisation tests for usage/lineage and cost-summary persistence.
tests/backend/test_subworkflow_budget_diagnosability.py Ensures persisted execution records include call IDs, usage, and cost summary fields.
tests/backend/test_settings_openai_model_probe.py Adds coverage for pre-client eligibility denial in OpenAI model probe route.
tests/backend/test_settings_llm_cost_summary.py Adds comprehensive tests for bounded historical cost summary + query bounding.
tests/backend/test_orchestrator_model_fallback_execution.py Verifies requested/selected/effective lineage propagation through fallback execution and heartbeat actor context.
tests/backend/test_openai_structured_tool_transport.py Ensures OpenAI transport captures cache token breakdown + service tier metadata.
tests/backend/test_model_registry_service.py Validates registry snapshot includes represented pricing JSON.
tests/backend/test_model_execution_eligibility.py New tests covering eligibility behaviour (local vs external, exact match, actorless failure).
tests/backend/test_llm_usage_cost_service.py New tests for usage normalisation, pricing, deduplication, partial/unavailable handling, and repricing.
tests/backend/test_llm_openai_temperature_guards.py Stubs eligibility for parameter transport tests to isolate concerns.
tests/backend/test_llm_api_key_resolution.py Stubs eligibility for legacy key resolution tests and adds actorless vision denial coverage.
tests/backend/test_adaptive_turn_service.py Adds tests for eligibility-denial terminal status and progress/cost reporting.
src/frontend/web/von_interface/templates/settings_tab.html Adds eligibility + cost summary UI elements and updates “allowed models” terminology.
src/frontend/web/von_interface/static/js/test/chatTab.test.js Adds Thinking UI tests for cost/usage summary and lineage rendering.
src/frontend/web/von_interface/static/js/settingsPage.js Implements eligibility-driven controls + historical cost-summary fetch/render logic.
src/frontend/web/von_interface/static/js/chatTab.js Renders per-turn usage/cost summary and per-call usage/lineage in Thinking.
src/backend/services/turn_execution_record_service.py Expands record normalisation to persist lineage + usage/cost summary from debug payloads.
src/backend/services/turn_execution_diagnostics_service.py Extends exchange normalisation and deduplication to keep call IDs, lineage, and usage.
src/backend/services/model_registry_service.py Adds pricing JSON predicate support and exposes pricing in model snapshot entries.
src/backend/services/llm_usage_cost_service.py New pure module to normalise usage and project estimated costs from represented pricing.
src/backend/services/llm_model_cost_history_service.py New bounded actor-scoped historical repricing/summary service for Settings.
src/backend/services/file_copy_interpretation_service.py Enforces eligibility gate for external vision providers and returns consistent denial shape.
src/backend/services/adaptive_turn_service.py Emits stable call IDs, captures provider-effective identity, and attaches cumulative cost summary to progress.
src/backend/server/routes/von_routes.py Computes/attaches per-turn llm_usage_cost_summary, threads registry snapshot into adaptive turns, and records support-call outcomes.
src/backend/server/routes/settings_routes.py Adds /api/settings/llm/cost_summary and enforces eligibility in OpenAI model probe route.
src/backend/server/routes/generate_route_support.py Records eligibility failures for presenter screen backfill and preserves llm_interaction in error debug info.
src/backend/languagemodels/structured_tool_calling/providers/openai_client.py Captures service tier and cached/cache-write token details from OpenAI responses.
src/backend/languagemodels/llm_interface.py Introduces ModelExecutionEligibilityError + assert_model_execution_allowed and wires checks into provider calls.
src/backend/integrations/internal_mcp/orchestrator.py Preserves contextvars across worker threads and refines requested/selected/effective lineage telemetry.
Review details
  • Files reviewed: 33/33 changed files
  • Comments generated: 1
  • Review effort level: Lite

We're testing this review assessment. Please use 👍 or 👎 to tell us if it's correct.

Comment on lines +142 to +146
return bool(
_model(model, provider=provider_key)
and _model(model, provider=provider_key)
== _model(candidate_model, provider=provider_key)
)
) -> None:
"""These tests isolate probe transport; policy denial has its own case."""

import src.backend.server.routes.settings_routes as settings_routes
def test_openai_model_probe_rejects_non_enabled_model_before_client_construction(
monkeypatch: pytest.MonkeyPatch,
) -> None:
import src.backend.server.routes.settings_routes as settings_routes
)
try:
cursor = cursor.sort("created_at_utc", -1).limit(safe_limit)
except AttributeError:
@witbrock
witbrock marked this pull request as ready for review August 5, 2026 12:37
@witbrock
witbrock merged commit ede52c2 into main Aug 5, 2026
5 checks passed
@witbrock
witbrock deleted the codex/JVNAUTOSCI-2624-premium-model-controls branch August 5, 2026 12:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants