Skip to content

vhpi: only arm the per-cycle phase callbacks once a plugin registers one - #1682

Open
djmazure wants to merge 1 commit into
nickg:masterfrom
djmazure:vhpi-lazy-phase-callbacks-v2
Open

djmazure wants to merge 1 commit into
nickg:masterfrom
djmazure:vhpi-lazy-phase-callbacks-v2

Conversation

@djmazure

@djmazure djmazure commented Oct 7, 2026

Copy link
Copy Markdown

Loading any VHPI plugin with --load roughly halves simulation speed, even when the plugin registers no per-cycle callbacks at all. vhpi_context_initialise arms the five phase callbacks (start/end of next cycle, end of process, last known delta cycle and end of time step) up front. They are then re-armed on every cycle, and each re-arm is a calloc/free pair, whether or not anything is registered for that reason.

This change arms those five reasons on demand from vhpi_register_cb, plus a sweep for callbacks registered before the context is initialised. The one-shot reasons are left armed as before. The phase is armed with the base reason rather than the vhpiCbRep* variant, so a repetitive callback keeps firing after the first time.

To reproduce, with this testbench and two small plugins:

library ieee;
use ieee.std_logic_1164.all;
use std.env.finish;

entity tb is
  generic (CYCLES : natural := 4000000);
end entity;

architecture test of tb is
  signal clk   : std_logic := '0';
  signal count : natural := 0;
begin
  clk <= not clk after 4 ns;

  process (clk) is
  begin
    if rising_edge(clk) then
      count <= count + 1;
      if count = CYCLES - 1 then
        finish;
      end if;
    end if;
  end process;
end architecture;

null.c registers nothing:

#include <vhpi_user.h>

static void startup(void)
{
}

void (*vhpi_startup_routines[])(void) = { startup, NULL };

rep.c counts vhpiCbRepEndOfTimeStep:

#include <stdio.h>
#include <vhpi_user.h>

static long fired;

static void on_step(const vhpiCbDataT *cb)
{
   fired++;
}

static void on_end(const vhpiCbDataT *cb)
{
   printf("vhpiCbRepEndOfTimeStep fired %ld times\n", fired);
}

static void startup(void)
{
   vhpiCbDataT step = { .reason = vhpiCbRepEndOfTimeStep, .cb_rtn = on_step };
   vhpiCbDataT end = { .reason = vhpiCbEndOfSimulation, .cb_rtn = on_end };
   vhpi_register_cb(&step, 0);
   vhpi_register_cb(&end, 0);
}

void (*vhpi_startup_routines[])(void) = { startup, NULL };
gcc -shared -fPIC -I<nvc>/src/vhpi -o null.so null.c
gcc -shared -fPIC -I<nvc>/src/vhpi -o rep.so rep.c
nvc -a tb.vhd -e tb
nvc -r tb
nvc -r --load ./null.so tb
nvc -r --load ./rep.so tb

Wall time for 4M clock cycles, best of 5, before and after builds interleaved (LLVM 14, Ubuntu 22.04):

before after
no plugin 0.969 s 0.969 s
--load null.so 2.122 s 1.003 s
--load rep.so 2.288 s 1.293 s

rep.so prints fired 7999999 times with both builds. Counting calloc calls with an LD_PRELOAD wrapper gives 48,002,514 before and 2,513 after with null.so, and 2,505 for both builds with no plugin, so this design was paying 12 allocations per clock cycle for a plugin that does nothing.

On a larger testbench I also saw runs with no plugin come out about 5% slower after the change, although that path executes none of the new code. A control build that adds the same functions without calling them showed most of that difference too, so I think it is code layout. The machine was not idle (load average 10 to 13), and the small testbench above shows no difference, but it may be worth a check on a quiet machine.

run_regr vhpi gives the same result before and after: 42 passed and 5 skipped, with Tcl not enabled in my build. I haven't added a test, as the existing vhpi2, vhpi13 and issue1505 tests cover the callback behaviour, but I'm happy to add one if you have a preferred way of testing this. I have only built and tested on Linux.

Loading any VHPI plugin currently costs about 40% of simulation throughput,
even a plugin that registers no callbacks at all.

vhpi_context_initialise() registers seven phase callbacks unconditionally as
soon as a plugin is --load'ed. Five of those reasons fire every delta cycle
(vhpiCbNextTimeStep, vhpiCbEndOfTimeStep, vhpiCbStartOfNextCycle,
vhpiCbLastKnownDeltaCycle, vhpiCbEndOfProcesses). The model frees each
callback node as it fires it, and vhpi_phase_cb immediately re-registers
itself, so model_set_phase_cb xcalloc()s a fresh node and walks the list —
every delta cycle, for callbacks nobody asked for. vhpi_run_callbacks then
iterates an empty array and returns.

Measured with an LD_PRELOAD allocation counter on a 4M-cycle testbench
(blinky + 7-segment mux + UART echo, 8 ns clock): a plugin registering
nothing makes 64,025,996 calloc calls, 16.006 per clock cycle, against
5,910 for a run with no plugin.

Arm the five per-cycle reasons only when a plugin actually registers a
callback for one, in vhpi_register_cb, and sweep any callbacks a startup
routine registered before the model existed. The one-shot reasons
(vhpiCbStartOfSimulation, vhpiCbEndOfSimulation) cost nothing per cycle and
stay armed unconditionally. Once armed, behaviour is unchanged.

The phase must be armed with the BASE reason, never the vhpiCbRep* spelling:
vhpi_run_callbacks derives its `rep` mapping by switching on the reason it is
handed, so a Rep reason arrives with rep == 0, matches cb->Reason == reason,
and the callback is consumed as a one-shot. An earlier revision of this patch
got that wrong and a plugin registering vhpiCbRepEndOfTimeStep fired exactly
once instead of every time step.

Results on the same testbench, 20M cycles, best of 7, before and after
runs interleaved on a shared machine:

  plain                1.930 -> 1.824 Mcyc/s
  + null plugin        1.130 -> 1.839
  --wave fst           0.875 -> 0.876
  --wave fst + plugin  0.655 -> 0.868   (+33%)

Allocations for the null plugin drop from 64,025,996 to 5,918, the same as a
run with no plugin. A plugin that does register vhpiCbRepEndOfTimeStep still
fires on every time step (8,000,000 times over 4M clock cycles) and pays one
allocation per firing (8,005,920 calloc calls).

`run_regr vhpi` gives 42 passed and 5 skipped (Tcl not enabled in this
build), identical to master, including vhpi2, vhpi13 and issue1505, which
register the repeating per-cycle callbacks this change defers.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant