Repository navigation
Conversation
Loading any VHPI plugin currently costs about 40% of simulation throughput, even a plugin that registers no callbacks at all. vhpi_context_initialise() registers seven phase callbacks unconditionally as soon as a plugin is --load'ed. Five of those reasons fire every delta cycle (vhpiCbNextTimeStep, vhpiCbEndOfTimeStep, vhpiCbStartOfNextCycle, vhpiCbLastKnownDeltaCycle, vhpiCbEndOfProcesses). The model frees each callback node as it fires it, and vhpi_phase_cb immediately re-registers itself, so model_set_phase_cb xcalloc()s a fresh node and walks the list — every delta cycle, for callbacks nobody asked for. vhpi_run_callbacks then iterates an empty array and returns. Measured with an LD_PRELOAD allocation counter on a 4M-cycle testbench (blinky + 7-segment mux + UART echo, 8 ns clock): a plugin registering nothing makes 64,025,996 calloc calls, 16.006 per clock cycle, against 5,910 for a run with no plugin. Arm the five per-cycle reasons only when a plugin actually registers a callback for one, in vhpi_register_cb, and sweep any callbacks a startup routine registered before the model existed. The one-shot reasons (vhpiCbStartOfSimulation, vhpiCbEndOfSimulation) cost nothing per cycle and stay armed unconditionally. Once armed, behaviour is unchanged. The phase must be armed with the BASE reason, never the vhpiCbRep* spelling: vhpi_run_callbacks derives its `rep` mapping by switching on the reason it is handed, so a Rep reason arrives with rep == 0, matches cb->Reason == reason, and the callback is consumed as a one-shot. An earlier revision of this patch got that wrong and a plugin registering vhpiCbRepEndOfTimeStep fired exactly once instead of every time step. Results on the same testbench, 20M cycles, best of 7, before and after runs interleaved on a shared machine: plain 1.930 -> 1.824 Mcyc/s + null plugin 1.130 -> 1.839 --wave fst 0.875 -> 0.876 --wave fst + plugin 0.655 -> 0.868 (+33%) Allocations for the null plugin drop from 64,025,996 to 5,918, the same as a run with no plugin. A plugin that does register vhpiCbRepEndOfTimeStep still fires on every time step (8,000,000 times over 4M clock cycles) and pays one allocation per firing (8,005,920 calloc calls). `run_regr vhpi` gives 42 passed and 5 skipped (Tcl not enabled in this build), identical to master, including vhpi2, vhpi13 and issue1505, which register the repeating per-cycle callbacks this change defers.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Loading any VHPI plugin with
--loadroughly halves simulation speed, even when the plugin registers no per-cycle callbacks at all.vhpi_context_initialisearms the five phase callbacks (start/end of next cycle, end of process, last known delta cycle and end of time step) up front. They are then re-armed on every cycle, and each re-arm is acalloc/freepair, whether or not anything is registered for that reason.This change arms those five reasons on demand from
vhpi_register_cb, plus a sweep for callbacks registered before the context is initialised. The one-shot reasons are left armed as before. The phase is armed with the base reason rather than thevhpiCbRep*variant, so a repetitive callback keeps firing after the first time.To reproduce, with this testbench and two small plugins:
null.cregisters nothing:rep.ccountsvhpiCbRepEndOfTimeStep:Wall time for 4M clock cycles, best of 5, before and after builds interleaved (LLVM 14, Ubuntu 22.04):
--load null.so--load rep.sorep.soprintsfired 7999999 timeswith both builds. Countingcalloccalls with anLD_PRELOADwrapper gives 48,002,514 before and 2,513 after withnull.so, and 2,505 for both builds with no plugin, so this design was paying 12 allocations per clock cycle for a plugin that does nothing.On a larger testbench I also saw runs with no plugin come out about 5% slower after the change, although that path executes none of the new code. A control build that adds the same functions without calling them showed most of that difference too, so I think it is code layout. The machine was not idle (load average 10 to 13), and the small testbench above shows no difference, but it may be worth a check on a quiet machine.
run_regr vhpigives the same result before and after: 42 passed and 5 skipped, with Tcl not enabled in my build. I haven't added a test, as the existing vhpi2, vhpi13 and issue1505 tests cover the callback behaviour, but I'm happy to add one if you have a preferred way of testing this. I have only built and tested on Linux.