Repository navigation
Implement internal prediction-distribution data producer - #414
Conversation
Implement `_prepare_probs_distribution_data` in `src/rtichoke/performance_data/probs_distribution.py` reproducing static binary prediction distribution statistical contract. Co-authored-by: uriahf <11351434+uriahf@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
|
|
This is close. The overall statistical contract, returned schemas, golden Please address the following focused issues before merge.
This is a correctness issue. The implementation currently does: However, That destroys the original input-shape information used by Consequently, input such as: is incorrectly interpreted as a model evaluated in the shared-population Build semantic metadata from the original Continue using Add an explicit regression test for exactly one keyed population confirming
Also retain the existing multiple-keyed-population test.
The implementation assigns interval indices once with This is still approximately O(n × number of intervals), which the task After calculating each observation’s interval index once, aggregate counts in Use either:
Polars is preferred because this package uses Polars and because a complete The intended pattern is approximately: Preserve:
A loop over evaluations is acceptable. A loop over operating-point rows for
The variable: is constructed but never used. Remove it, or use one single complete interval-grid object as the authoritative
Please finish the naming cleanup in the production module. Rename, where applicable: Short mathematical names inside tiny tests are acceptable. This request applies
This code: cannot currently detect anything because Python dictionary keys are unique by The meaningful invariant is uniqueness of: across the metadata values. Validate that directly: and reject duplicates there. If the current public dictionary input makes duplicates practically
Add a test with:
The test need not assert a fragile wall-clock duration. Its purpose is to If practical, isolate interval aggregation in a small private helper whose code
Do not change:
After these changes, run and report: Please report the updated head SHA and confirm that the single keyed-population |
Thank you for the clear, targeted feedback! I have implemented all requested changes:
|
…ucer - Preserve single keyed-population metadata semantics by passing unaligned reals to `_build_evaluation_metadata`. - Vectorize interval aggregation via Polars `group_by(\"interval_id\")` and left join onto a complete interval grid DataFrame in `_aggregate_bins_for_evaluation`. - Remove unused interval list construction. - Clean up production variable names and validate duplicate `metadata.evaluation`. - Add performance regression test and single keyed-population test. Co-authored-by: uriahf <11351434+uriahf@users.noreply.github.com>
|
Thank you — I re-reviewed the updated implementation at head The substantive issues from the previous review are resolved:
The remaining small loops operate over interval definitions or evaluation One minor documentation issue: the PR description still reports the original Approved from the statistical-contract and cross-language architecture |
Thank you for the thorough review and approval! I have updated the submission description to reflect the final head SHA. |
Deliverable Summary
93fe1e8b4380e18ccd09849af429bb8baf3ddadd(v0.1.36)1ce587424fb8e9c647565406085a85ae763ef4cbsrc/rtichoke/performance_data/probs_distribution.pytests/test_probs_distribution.py_prepare_probs_distribution_datainsrc/rtichoke/performance_data/probs_distribution.py_validate_and_align_binary_inputsfromrtichoke.performance_data.performance_data_build_evaluation_metadatafromrtichoke.processing.evaluation_semanticsbins:evaluation(pl.String),model(pl.String, nullable),population(pl.String),lower(pl.Float64),upper(pl.Float64),include_lower(pl.Boolean),include_upper(pl.Boolean),n_positive(pl.UInt32),n_negative(pl.UInt32)operating_points:evaluation(pl.String),model(pl.String, nullable),population(pl.String),type(pl.String),value(pl.Float64),cutoff(pl.Float64),realized_ppcr(pl.Float64)TP=total_pos, FP=total_neg, TN=0, FN=0).probabilities <= cutoffon predicted-negative side.[0.0, 0.0]is always present for zero-score observations, followed by right-closed intervals(lower, upper].value, effectivecutoff, andrealized_ppcrare preserved distinctly.binsmatchprepare_performance_data()output across all operating points and scenarios.uv run pytest(349 passed, 9 skipped)uv run ruff check .uv run ruff format --check .uv run ty check src/rtichoke/performance_data/probs_distribution.py tests/test_probs_distribution.pyPR created automatically by Jules for task 7515411089152951279 started by @uriahf