Require explicit premium model eligibility and surface turn cost - #348
Merged
Merged
Conversation
| }); | ||
|
|
||
| test('uses the persisted effective allow list, not staged edits', () => { | ||
| const eligibility = document.getElementById('openaiModelEligibilityStatus'); |
There was a problem hiding this comment.
🔵 Human review recommended
It changes core model-execution authority enforcement and cost/usage telemetry across multiple backend and frontend paths, warranting final human review despite strong test coverage.
Pull request overview
Implements scoped, explicit eligibility checks for running external/premium generative models (reusing enabled_llms as the execution authority) and adds content-free token/cost evidence so the UI can surface per-turn estimated costs (Thinking) plus bounded recent historical cost summaries (Settings).
Changes:
- Add a shared backend eligibility gate (
assert_model_execution_allowed) and wire it into structured tool calls, legacy provider calls, settings probes, and vision/image paths. - Persist and normalise richer per-call model identity lineage (requested/selected/effective + source), token usage (including cache breakdown), and compute a per-turn
llm_usage_cost_summary. - Add an actor-scoped historical cost summary service + Settings endpoint and render eligibility/cost status in the Settings UI and Thinking card.
File summaries
| File | Description |
|---|---|
| tests/frontend/settingsOpenAiModelChange.test.js | Updates Settings page expectations for eligibility/cost UI and “allowed” wording. |
| tests/frontend/settingsModelPoolControls.test.js | Adds tests for persisted allow-list precedence and OpenAI cost-summary fetching/rendering. |
| tests/backend/test_von_turn_execution_debug_info.py | Validates finalised debug payload includes cost summary and preserves paid calls on failure. |
| tests/backend/test_von_generate_workflow_instances.py | Ensures model registry snapshot is passed through generate/adaptive turn and eligibility denials are recorded. |
| tests/backend/test_thinking_llm_call_timestamps_and_precedence.py | Extends exchange normalisation tests for usage/lineage and cost-summary persistence. |
| tests/backend/test_subworkflow_budget_diagnosability.py | Ensures persisted execution records include call IDs, usage, and cost summary fields. |
| tests/backend/test_settings_openai_model_probe.py | Adds coverage for pre-client eligibility denial in OpenAI model probe route. |
| tests/backend/test_settings_llm_cost_summary.py | Adds comprehensive tests for bounded historical cost summary + query bounding. |
| tests/backend/test_orchestrator_model_fallback_execution.py | Verifies requested/selected/effective lineage propagation through fallback execution and heartbeat actor context. |
| tests/backend/test_openai_structured_tool_transport.py | Ensures OpenAI transport captures cache token breakdown + service tier metadata. |
| tests/backend/test_model_registry_service.py | Validates registry snapshot includes represented pricing JSON. |
| tests/backend/test_model_execution_eligibility.py | New tests covering eligibility behaviour (local vs external, exact match, actorless failure). |
| tests/backend/test_llm_usage_cost_service.py | New tests for usage normalisation, pricing, deduplication, partial/unavailable handling, and repricing. |
| tests/backend/test_llm_openai_temperature_guards.py | Stubs eligibility for parameter transport tests to isolate concerns. |
| tests/backend/test_llm_api_key_resolution.py | Stubs eligibility for legacy key resolution tests and adds actorless vision denial coverage. |
| tests/backend/test_adaptive_turn_service.py | Adds tests for eligibility-denial terminal status and progress/cost reporting. |
| src/frontend/web/von_interface/templates/settings_tab.html | Adds eligibility + cost summary UI elements and updates “allowed models” terminology. |
| src/frontend/web/von_interface/static/js/test/chatTab.test.js | Adds Thinking UI tests for cost/usage summary and lineage rendering. |
| src/frontend/web/von_interface/static/js/settingsPage.js | Implements eligibility-driven controls + historical cost-summary fetch/render logic. |
| src/frontend/web/von_interface/static/js/chatTab.js | Renders per-turn usage/cost summary and per-call usage/lineage in Thinking. |
| src/backend/services/turn_execution_record_service.py | Expands record normalisation to persist lineage + usage/cost summary from debug payloads. |
| src/backend/services/turn_execution_diagnostics_service.py | Extends exchange normalisation and deduplication to keep call IDs, lineage, and usage. |
| src/backend/services/model_registry_service.py | Adds pricing JSON predicate support and exposes pricing in model snapshot entries. |
| src/backend/services/llm_usage_cost_service.py | New pure module to normalise usage and project estimated costs from represented pricing. |
| src/backend/services/llm_model_cost_history_service.py | New bounded actor-scoped historical repricing/summary service for Settings. |
| src/backend/services/file_copy_interpretation_service.py | Enforces eligibility gate for external vision providers and returns consistent denial shape. |
| src/backend/services/adaptive_turn_service.py | Emits stable call IDs, captures provider-effective identity, and attaches cumulative cost summary to progress. |
| src/backend/server/routes/von_routes.py | Computes/attaches per-turn llm_usage_cost_summary, threads registry snapshot into adaptive turns, and records support-call outcomes. |
| src/backend/server/routes/settings_routes.py | Adds /api/settings/llm/cost_summary and enforces eligibility in OpenAI model probe route. |
| src/backend/server/routes/generate_route_support.py | Records eligibility failures for presenter screen backfill and preserves llm_interaction in error debug info. |
| src/backend/languagemodels/structured_tool_calling/providers/openai_client.py | Captures service tier and cached/cache-write token details from OpenAI responses. |
| src/backend/languagemodels/llm_interface.py | Introduces ModelExecutionEligibilityError + assert_model_execution_allowed and wires checks into provider calls. |
| src/backend/integrations/internal_mcp/orchestrator.py | Preserves contextvars across worker threads and refines requested/selected/effective lineage telemetry. |
Review details
- Files reviewed: 33/33 changed files
- Comments generated: 1
- Review effort level: Lite
We're testing this review assessment. Please use 👍 or 👎 to tell us if it's correct.
Comment on lines
+142
to
+146
| return bool( | ||
| _model(model, provider=provider_key) | ||
| and _model(model, provider=provider_key) | ||
| == _model(candidate_model, provider=provider_key) | ||
| ) |
| ) -> None: | ||
| """These tests isolate probe transport; policy denial has its own case.""" | ||
|
|
||
| import src.backend.server.routes.settings_routes as settings_routes |
| def test_openai_model_probe_rejects_non_enabled_model_before_client_construction( | ||
| monkeypatch: pytest.MonkeyPatch, | ||
| ) -> None: | ||
| import src.backend.server.routes.settings_routes as settings_routes |
| ) | ||
| try: | ||
| cursor = cursor.sort("created_at_utc", -1).limit(safe_limit) | ||
| except AttributeError: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Outcome
Implements JVNAUTOSCI-2624 so external generative models can run only when the exact provider/model is in the existing actor- and organisation-scoped allowed-model pool. Persisted, content-free provider usage and effective-model evidence now supports honest per-turn cost estimates in Thinking and bounded recent-cost summaries in Settings.
Smallest coherent design
enabled_llmsas the single scoped execution authority; enforce it at the shared generation boundary and the few raw vision/live-probe call sites.Embeddings are intentionally outside chat-model eligibility: they are fixed non-generative infrastructure, while the RAG runtime LLM already passes through the gated generation boundary. Adding embeddings to
enabled_llmswould expand this task and break the existing retrieval contract.Represented pricing evidence
#V#has_model_pricing_json:40543e7b-e706-40db-8df4-b58608456a6e6a732a105c780564a20705226a732a265c780564a2070523openai-standard-observed-2026-08-05Validation
ruff,py_compile, and diff checks passedThe existing repository-level Sol prohibition remains unchanged; this change neither invokes nor restores Sol.