Skip to content

Rust Rewrite - #375

Draft
sandorstormen wants to merge 49 commits into
mainfrom
rust-rewrite
Draft

sandorstormen wants to merge 49 commits into
mainfrom
rust-rewrite

Conversation

@sandorstormen

Copy link
Copy Markdown
Contributor

No description provided.

sandorstormen and others added 30 commits October 4, 2026 06:31
Replace hardcoded log tuples tied to a specific RNG seed with forced
detector tprate/fprate extremes (1.0/0.0/-1.0) and invariant-based
assertions, so the tests no longer break if simulator internals change
how/when rng draws happen. Also fixes detector_lang.mal's detector
annotation syntax (fnr -> tpr) to match the current mal-toolbox grammar.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A dyna-MAL model effect can delete the asset performing the step
(e.g. `R> sub*.files / self` with no `~` marker, which deliberately
means "destroy self", not just unlink an association). Every attack
step on that asset, including the step itself, is then removed from
the live attack graph.

mal-toolbox's core always keeps a *live* node's children/parents
correctly unlinked from anything removed, but a node that is itself
dead only exposes a frozen pre-removal `.children` snapshot (kept
readable so held references, e.g. in `performed_nodes`, don't crash).
get_attack_surface/get_effects_of_attack_step were reading that stale
snapshot as if it were live, letting an already-removed node leak back
into a later action surface and crash several steps after the fact
with an unrelated-looking KeyError.

Skip expanding `.children` from any node that is no longer live - the
single point where a stale snapshot could be consulted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Relocates the malsim package to python/malsim so it sits alongside new
Rust code, mirroring mal-toolbox's rust-rewrite layout. Adds a core/
Cargo workspace (malsim-core, depending on mal-toolbox's
maltoolbox-attackgraph crate on the rust-rewrite branch) and a separate
py-bindings/ Cargo workspace (malsim-pyo3) for PyO3 bindings between
the Rust core and the Python package.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Documents the phased plan (MalSimulator, then DynaMalSimulator, then a
Rust-only library/scenario-loading path) for moving the hot step()/reset()
loop into core/malsim-core + py-bindings/malsim-pyo3, while keeping
AttackerState/DefenderState/Scenario/config objects as unchanged,
hand-written Python classes. Records the architectural decisions made
along the way (RNG reproducibility, shared-graph strategy with
mal-toolbox's PyO3 layer, NodePropertyRule/reward boundary, Rust code
style and terminology) and a running differences log to track Rust-vs-
Python divergences as they're found.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…l-toolbox PyO3 modules

Switches pyproject.toml to maturin and wires malsim-pyo3 to mal-toolbox's
rust-rewrite branch (pinned to commit c854d1d6) to validate PORTING_NOTES.md
§2.2 end to end. The originally-specified mechanism (downcasting a Python
maltoolbox.AttackGraph back to mal-toolbox's PyAttackGraph pyclass from
inside a separately-compiled cdylib) doesn't work - PyO3 pyclasses from a
shared dependency crate get an unrelated type object in every cdylib that
statically links it. Fixed via a PyCapsule handoff (PyAttackGraph's new
__inner_capsule__() upstream, consumed here), which also drops malsim-pyo3's
dependency on maltoolbox-attackgraph-py entirely. Also fixes a double-free
in the capsule consumer found via a create/call/delete/gc.collect() stress
test (needed Rc::increment_strong_count before Rc::from_raw).

See PORTING_NOTES.md §10 for the full writeup.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ports ttc_utils.py's DistFunction/Operation/TTCDist (expected_value,
sample_value, success_probability, attempt_ttc_with_effort,
attempt_bernoulli, from_dict/to_dict) to core/malsim-core/src/ttc.rs,
backed by the statrs crate for CDF/mean/sampling. Deliberately excludes
the graph-node-dependent default_ttc_dist/from_node/from_name, scoped to
Phase A3 instead, so this module stays independent of
maltoolbox_attackgraph.

26 Rust-native tests cover every DistFunction variant's expected_value
(exact, since it's closed-form) plus the combine_with/combine_op
composition suite ported 1:1 from test_ttc_utils.py, and
structural/statistical (not seed-pinned) checks for the RNG-touching
methods per PORTING_NOTES.md's RNG-equivalence policy.

PORTING_NOTES.md updated: A2 marked done, a seed-pinned-test follow-up
noted for A9, and a differences-log entry for the statrs parameter-order
gotchas (Binomial, Gamma) and the deliberately-preserved
success_probability/combine_with behavior.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ports graph_state.py and the graph-node-dependent half of ttc_utils.py
(default_ttc_dist, TTCDist.from_node, attack_step_ttc_value(s),
get_pre_enabled_defenses, get_impossible_attack_steps,
compute_initial_graph_state) into core/malsim-core/src/graph_state.rs,
and graph_processing.py's necessity propagation (evaluate_necessity,
_propagate_necessity_from_node, calculate_necessity) into
core/malsim-core/src/necessity.rs. Viability is deliberately not
ported - confirmed dead/deprecated code with no callers outside its
own module.

Each graph-dependent function is split into a thin AttackGraphNode-
reading wrapper plus a graph-independent helper over plain data, so
the actual logic stays unit-testable even without a real attack graph
fixture (building one needs a dev-dependency on maltoolbox-language,
deferred per discussion - see PORTING_NOTES.md §10). 41 Rust tests
passing; cargo clippy/fmt clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Ports node_is_blocked/node_is_traversable/node_is_live/is_attack_step
from graph_utils.py into core/malsim-core/src/graph_utils.rs
(node_is_actionable/node_reward stay Python-only per §2.4).

Adds maltoolbox-language as a dev-dependency-only crate dependency so
tests can build real AttackGraphNode fixtures (core/malsim-core/src/
test_fixtures.rs, mirroring tests/conftest.py::dummy_lang_graph) -
also backfills necessity.rs's previously-deferred tests from A3 in the
same change. 59 Rust-native tests total, cargo clippy/fmt clean; full
Python suite still green.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…lsim-core)

Ports get_attack_surface, get_effects_of_attack_step (attack_surface.py)
and get_defense_surface (defense_surface.py). Actionability is taken as
an already-flattened id-set rather than a NodePropertyRule, per the
porting plan. 14 new Rust-native tests; cargo test/clippy/fmt and the
full Python suite stay green.
…core)

Ports false_alerts.py/observability.py/event_logger.py into
core/malsim-core/src/{false_alerts,observability,event_logger}.rs: rate
rules are taken as already-flattened id maps/sets per §2.4 (same pattern
as A5), and LogEntry's detector/trigger fields become id-based since
mal-toolbox's Detector has no id of its own.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…alsim-core)

Composes A2-A6's modules into the actual per-agent step logic: attacker_step/
attempt_attacker_step/attacker_step_effects/attacker_is_terminated and
defender_step/defender_is_terminated, with 22 new Rust-native tests. See
PORTING_NOTES.md §0/§10 for scope and the preserved-as-is behavioral quirks
this phase ported faithfully.
…ive)

Composes A1-A7's ported malsim-core modules into a full reset/step loop
exposed to Python as plain dicts, proving they work together end to end
before A9 wires MalSimulator itself to the native backend.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
MalSimulator now holds one malsim._native.Simulator instance and
rebuilds AttackerState/DefenderState/MalSimulatorState from its
reset_native()/step_native() output via new native-driven factory
functions, added alongside (not replacing) the pure-Python ones
DynaMalSimulator still depends on.

Extends the native layer to close A8's remaining scope cuts: per-agent
ttc_dists overrides are now parsed and applied in reset_native, and
reset_native/step_native expose the full GraphState (ttc_values,
impossible_attack_steps, necessity_per_node, pre_enabled_defenses) so
Python can reconstruct MalSimulatorState without recomputing anything.

Fixes two bugs surfaced by this phase's first real end-to-end runs:
- A pre-existing collect_logs bug (checked model_asset even for nodes
  with no detectors; Python never did).
- A genuine cross-process non-determinism bug: Rust's default HashSet
  hasher is randomized per process, so action_surface/other id lists
  crossing the FFI boundary came back in a different order on every
  run even for an identical seed - stable_ids now sorts.

Also documents and works around two mal-toolbox PyO3 binding quirks
(node.children/.parents and node.detectors are Python-side caches that
.add()/item-assignment never sync back to the real graph) that broke
~35 test call sites and 3 detector-rate tests once native started
reading the graph directly instead of through Python.

Full details, including the DynaMalSimulator-coupling rationale for
the new-functions-alongside-old-ones design and every relaxed/re-pinned
seed-dependent test, are in PORTING_NOTES.md §0/§9/§10.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
Adds a rust.yml workflow (fmt + build, matrixed over the root and
py-bindings cargo workspaces) and switches the PyPI publish workflow to
maturin-action's multi-platform wheel matrix with TestPyPI gating PyPI,
matching mal-lang/mal-toolbox#255.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
The linux x86_64 "Test wheel" step reinstalls mal-simulator from the
freshly built wheel, which pulls in mal-toolbox from git and compiles
it from source on the host. maturin-action's prior sccache-enabled
build step leaves RUSTC_WRAPPER=sccache in the job environment, but
that binary only exists inside the manylinux container, so the host
build failed with "could not execute process `sccache ... rustc -vV`".

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
Confirms all 19 audited public query methods already satisfy the
"native-backed vs. pure-Python mirror" split A9 established, with zero
code changes needed - verified via git diff across the A8->A9 boundary.
Also flags a new drift risk (§9) for node_is_blocked/node_is_traversable's
duplicate Python/Rust implementations and clarifies graph_utils.py/
state_query.py/node_getters.py are not A11 deletion candidates.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
…ated state

Profiling showed the native rewrite was 2.6x slower than pre-port pure
Python (111.6s vs 42.7s for 5000 steps): step_native resent the entire
episode-accumulated state over the FFI boundary every call even though
it had already computed the per-step delta and discarded it, and both
Python state factories re-resolved/re-parsed the full value from scratch
every step instead of merging against previous_state. The defender logs
field was the worst case - O(episode^2), since every log ever fired was
re-parsed into a fresh LogEntry on every single step.

step_native's output is now delta-only for monotonically-growing fields
(step_enabled_defenses/step_performed_nodes/step_attempted_nodes/
step_compromised_nodes/step_observed_nodes/step_logs); episode-static
fields (ttc_values, necessity_per_node, impossible_attack_steps,
pre_enabled_defenses) are dropped from step_native's output and cached
from reset_native instead. Both *_state_factories.py now merge deltas
against previous_state. reset_native's output shape is unchanged.

Re-profiling confirms the fix: a 5000-step run accumulating 10,002 logs
(worst case for the old O(episode^2) behavior) completes in ~1s with
flat per-step cost.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Break §6 into B1-B7: 2 pure-Rust-logic steps in malsim-core (compressed
from Phase A's 7, since mal-toolbox's rust-rewrite branch already ports
the hard primitives - model mutation and the model-effect grammar) plus
5 Python-bindings-integration steps mirroring Phase A's A1/A8-A11
granularity, including a new B3 proving step for a shared Model handle
that Phase A never needed (PyModel has no __inner_capsule__ yet).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
… step orchestration to malsim-core

B1 ports association-traversal evaluation + model-effect application
(process_assoc_traversal.py/model_effects.py) to assoc_traversal.rs/
model_effects.rs, operating on i64 asset ids throughout rather than
object references. B2 composes B1 with Phase A's existing
attacker_step/defender_step/graph_state functions into the full dyna
step orchestration (dyna_graph_state.rs/dyna_attacker_step.rs/
dyna_defender_step.rs) plus model-snapshot reconciliation for reset
(model_state.rs).

Promotes maltoolbox-model and maltoolbox-language to direct malsim-core
dependencies. 23 new Rust-native tests (149 total), reusing wiperLang.mal/
wiper_model.yml fixtures already used by the Python test suite. No Python
files touched; full Python/mypy/ruff gate stays green.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
Mirrors A1's PyAttackGraph capsule approach for maltoolbox.Model.
Upstream mal-toolbox gained PyModel::__inner_capsule__ (commit
b96258bad474282b975245d9848f7c25d195d508 on rust-rewrite); this repo
re-pins to that rev, adds maltoolbox-model as a direct malsim-pyo3
dependency, and wires up extract_shared_model plus two test-support
pyfunctions (model_asset_count, model_add_asset_native) that prove
mutations are visible from both the Python and native sides - the one
thing A1's own smoke test never needed to prove, since Phase A never
mutates the shared graph from both sides at once.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
B4 (done): extends the existing malsim._native.Simulator pyclass with
dyna_reset_native/dyna_step_native (reusing AttackerRuntime/DefenderRuntime/
SimState and the output-building helpers rather than a separate pyclass),
threading B3's shared Model handle alongside A1's shared AttackGraph handle.
The pristine model snapshot is captured natively on first use instead of
being passed in from Python. malsim-core's AssetOp/AssocOp now carry a
self-contained AssetRef (id+type+name) per asset reference, captured at
record time, so a modification record never needs a live Model lookup to
resolve - needed because a later op in the same record can invalidate an
id an earlier op referenced, and native-driven asset removal bypasses
maltoolbox's PyO3-level tombstone mechanism entirely.

B5 (partial, blocked): DynaMalSimulator.reset()/.step() are rewired to
delegate to native, mirroring Phase A9's reset()/step() exactly. Blocked
on a real regression: native dyna stepping can panic the whole process
("invalid SlotMap key used") when a model effect's association churn
disconnects an asset, since several Phase-A-era malsim-core functions
(get_attack_surface, attempt_attacker_step, necessity::calculate_necessity,
node_is_blocked/node_is_traversable, collect_logs) do unguarded node-id
indexing that assumed the graph never mutates mid-episode - an assumption
Phase B's whole design breaks. Full findings, a repro, and three
undecided remediation options are recorded in PORTING_NOTES.md's B5
entry and in section 9's open risks for whoever resumes this phase.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
…TTC gap

Hands-on eprintln instrumentation (not just backtrace reading) found the
B5 panic's real mechanism differs from the original PORTING_NOTES writeup:
a step's own model effect can remove the step's own backing asset outright
(no regeneration, no new id), not "regenerate a disconnected asset's node
under a new id" as previously diagnosed - this is the second time this
failure class was misdiagnosed, now called out explicitly in the notes.

Guard every unguarded graph.nodes[id] indexing on accumulated/historical
ids found via a full audit of both crates: collect_logs (the original
crash site), dyna_attacker_step's effect-chain loop, and - not in the
original 5-site list at all - malsim-pyo3::simulator.rs's stable_ids/
id_value_map/attacker_ttc_overrides/log_entry_to_py, where the second
domino crash actually was. test_different_attackers: 0/12 -> 10/12.

Remaining 2/12 (TTCSoftMinAttacker) fail on a different, newly-exposed
bug - newly-created dyna nodes missing a TTC value - confirmed as a
genuine Rust-port regression via comparison against pure-Python
DynaMalSimulator, documented with a repro for whoever resumes B5.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
…the model

The dyna native step was treating ttc_values/necessity_per_node/
impossible_attack_steps/pre_enabled_defenses as episode-static (true for
Phase A, false once model effects grow them mid-episode), so a node
created by a dyna model effect never got its ttc value mirrored back to
Python. build_step_output now resends the full current maps, but only on
a step whose step_modification_record is non-empty - correctness over a
per-field delta, since necessity_per_node is a whole-graph recompute, not
just new keys.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
A long-standing typo (dot became a comma) left this as a non-.py file.
It never broke anything until maturin packaged it into a wheel RECORD,
where the unescaped comma corrupted the CSV and broke `uv sync` for any
project depending on mal-simulator via git.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0116xUQqCURws4Hk4w4cBdKo
claude and others added 19 commits October 9, 2026 13:33
…ery regression test

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WQHkC7DzWcwzFL8tJM5jH
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WQHkC7DzWcwzFL8tJM5jH
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015WQHkC7DzWcwzFL8tJM5jH
…to Rust

Deletes the Python hot-path code that MalSimulator/DynaMalSimulator no
longer reach now that reset/step run natively: attack_surface.py,
defense_surface.py, graph_processing.py, reset_agent.py,
simulator_static_data.py and dyna's attacker_step/defender_step/
graph_state/model_effects/model_state/process_assoc_traversal.py, plus
the dead functions in event_logger/false_alerts/observability/ttc_utils/
node_getters/graph_utils/*_state_factories. Live items stay in their
original modules (LogEntry, GraphState, *_is_terminated).

graph_processing.py's viability/pruning half had no Rust port; it is
ported to malsim-core (viability.rs) with its four Python tests so their
coverage survives the deletion.

Python tests of deleted code are removed where Rust tests cover them;
test_attack_surface_traininglang and the dyna partial-regeneration test
are rewritten to assert the same properties through the live simulator
API instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019PMhGDtMmmb2JEKC2M8Zco
…_NOTES

New Rust tests close the gaps found while checking that every deleted
Python test is covered: get_attack_surface with enabled defenses and
impossible steps, an attacker action on a defense node, a defender
action outside a non-empty surface, and node_is_traversable returning
false behind an enabled defense. model_effects.rs's graph-equivalence
helper now compares edges by full name and checks node count, and the
removal test checks for dangling associations.

PORTING_NOTES: A11, B7, Phase A and Phase B marked done. New
Conventions (§11) and Decisions (§12) sections. A §9 risk for native
memory growth across repeated dyna runs: it predates the port and
valgrind attributes it to mal-toolbox, not malsim's capsule handoff.
§10 records the viability port's deviations and the dyna settings
snapshot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019PMhGDtMmmb2JEKC2M8Zco
…aming, rule typing, parity mechanism)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…o module)

Independent Rust NodePropertyRule<T: RuleValue> for the Rust-only scenario
path: by_asset_name > by_asset_type > default precedence with Python's
truthiness fall-through, list vs mapping step forms, per_node by node id,
and RuleValue impls for bool/f64/TtcDist.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
… resolution) to malsim-core

load_scenario_dict/recursive_update/path_relative_to_file_dir/
validate_scenario_dict over serde_json::Value, parsed with serde_yaml.
Tests cover merge semantics, validation, and every fixture's validate
outcome matching Python's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…ngs module conventions

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…im-core

New settings module (MalSimulatorSettings, AttackSurfaceSettings,
RewardMode, Flat*Settings simulator inputs); scenario::agent_settings
(AttackerSettings<N>/DefenderSettings/agent_settings_from_dict, policy
kept opaque); scenario::flatten (get_entry_points + flatten_*_settings
mirroring native_settings.py's per-node semantics).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…-Rust graph construction

Scenario::{new, from_dict, load_from_file} loads the language and model
(file or inline), builds the AttackGraph, resolves attacker entry
points/goals, and exposes flatten_agents() producing Simulator::reset
input. Ports test_scenario.py's scenario-level cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…m-pyo3 into malsim-core

New public malsim_core::simulator::Simulator (new/new_dyna/reset/step/
state) with SimState, Attacker/DefenderRuntime, StepOutcome and
SimulatorError. malsim-pyo3 is now a thin dict-parsing/output-building
wrapper with an unchanged Python-facing API. Agents are iterated in name
order (deterministic) instead of HashMap order.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…Scenario -> Simulator -> steps)

Integration test in malsim-core using only the public API: loads real
fixtures, builds a Simulator (plain and dyna), resets via
Scenario::flatten_agents and steps with hand-picked actions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…n JSON

tests/scenario_parity.py generates each fixture's resolved shape from the
Python Scenario into tests/testdata/scenario_parity.json;
tests/test_scenario_parity.py keeps it current, and
core/malsim-core/tests/scenario_parity.rs checks the Rust loader against
it. Fixes the gap it found: YAML !!set entry points/goals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…'s graph

DynaMalSimulator.reset flattened per-node rules (actionable/observable
steps, FP/FN rates, TTC overrides) before native restored the model, so
nodes removed by the previous episode's model effects and regenerated by
the restore were missing from the new episode's settings.

Adds the additive core Simulator::restore_model() (no-op for plain
simulators and pristine models; reset still restores), exposed as pyo3's
dyna_restore_model_native(model). dyna_reset now restores first, then
flattens. Regression tests in Python and Rust (both fail without the
fix); Scenario::flatten_agents documents the restore-first requirement.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
…restore

Entry points/goals were node objects (Python) / node ids (Rust) resolved
once at scenario load. When an episode's model effects removed their
asset, the restore regenerated those nodes with new ids: the next Python
dyna reset raised 'node id N is not part of this simulator's attack
graph', and the Rust-only path silently dropped the entry point.

DynaMalSimulator now captures full-name attacker settings at construction
and dyna_reset re-resolves them after dyna_restore_model_native, returning
the refreshed agent settings. The Rust Scenario keeps a private full-name
copy and flatten_agents re-resolves it on every call (now returns
Result). Regression tests in Python and Rust for single and multiple
entry-point sets; both fail without the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MjG65UvNdYrV9uGiqo8nXa
Split malsim-core's single Simulator (new_dyna + optional model handle)
into two types: Simulator runs a static instance model/attack graph, and
the new DynaSimulator wraps a Simulator plus the shared Model and its
snapshot, reusing the Simulator's reset and per-agent bookkeeping and
swapping in the model-effect-aware step functions (composition, the Rust
analogue of DynaMalSimulator(MalSimulator)). restore_model moves to
DynaSimulator, and StepOutcome drops the modification record, which only
DynaStepOutcome carries.

malsim-pyo3 exposes the same split as two pyclasses:
_native.Simulator(graph) and _native.DynaSimulator(graph, model), both
with reset_native/step_native (plus restore_model_native on the dyna
one), sharing the dict conversions. The dyna_*_native methods are gone.
DynaMalSimulator builds a DynaSimulator in __init__ and keeps it as
_dyna_native_sim.

Record the user's decisions in PORTING_NOTES.md §11 (superseding the
single-Simulator convention) and the judgment calls in §12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017CNWy9A9w7qn4m6YfR1nDY

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants