Skip to content

Speed up the inner interpreter - #60

Open
jameshickman wants to merge 2 commits into
zevv:masterfrom
jameshickman:perf/inner-loop
Open

jameshickman wants to merge 2 commits into
zevv:masterfrom
jameshickman:perf/inner-loop

Conversation

@jameshickman

Copy link
Copy Markdown

Two changes to the inner interpreter, worth about 1.6x at -O2, plus a build flag change worth another 2.1x on top of that. No behaviour change: output is byte identical on every script in forth/, including -t trace output.

Method

A 5M iteration Forth loop, instrumented by patching a counter into run() so the dispatch count is exact rather than estimated: 65,002,186 dispatches (60,002,041 primitives + 5,000,145 word calls). All figures are best of five runs on one machine, counters from perf stat. The zf_cell type, dictionary and stack sizes are the stock src/linux/zfconf.h.

What the profile actually says

The interesting part is what is not the problem. On the current master at -O2:

2.45e9 cycles   7.44e9 instructions   IPC 3.03
branch-misses:      14,970  /  2.05e9 branches   (0.0007%)
L1-dcache-misses:   24,623  /  2.48e9 loads      (0.001%)

IPC of 3.0 with effectively zero misses means the CPU is never stalled. It is running flat out, executing 114 x86 instructions per bytecode dispatch. This is an instruction count problem, not a dispatch, prediction or memory problem.

In particular the switch(op) in do_prim() is already fine. zf_prim is dense and contiguous, so GCC emits a jump table, do_prim() is static with one call site and gets fully inlined into run(), and perf annotate puts the resulting jmp *0x...(,%rdx,8) at 0.00% of samples. There is nothing to win there. The time is going into work done per instruction, so both changes below remove some of that work.

1. Decode single byte opcodes inline in run()

perf report on the -O2 build with tracing compiled out:

76.8%  run
22.0%  dict_get_cell_typed

Every instruction fetch made an out of line call into dict_get_cell_typed(), plus a second one per lit operand, to decode what is usually a single byte. Most dictionary cells encode as one byte, so that case is now decoded inline and the call is kept only for the multi byte path:

uint8_t b0 = ctx->dict[ctx->ip];
if((b0 & 0x80) == 0) { d = b0; code = b0; l = 1; }
else                 { l = dict_get_cell(ctx, ctx->ip, &d); code = d; }

Worth 33% on its own.

2. Only walk the return stack for the trace indent when tracing is on

for(i=0; i<RSP(ctx); i++) trace(ctx, "┊  ");

ran once per instruction whether or not tracing was enabled. trace() already guards its body, so nothing was printed, but the loop still evaluated RSP(ctx) every iteration. That is ctx->uservar[ZF_USERVAR_RSP], and since ctx->uservar is set to (zf_addr *)ctx->dict it aliases the dictionary, so the compiler cannot keep it in a register and has to reload it through the pointer after every dictionary or stack write. Wrapping the loop in a single if(TRACE(ctx)) removes it from the non-tracing path.

This is the larger of the two. It also explains a result that surprised me: compiling tracing in costs 45-50% of run time even when tracing is switched off at run time, because every trace site still has to perform that aliased load and test.

Results

Build Time Instr/dispatch IPC
master, -Os (what the Makefile does today) 1.24s 158 1.99
master, -O2 0.58s 114 3.06
master, -O2, tracing compiled out 0.29s 64 3.43
this PR, -O2 0.37s 80 3.32
this PR, -O2, tracing compiled out 0.20s 44 3.24

1.57x at -O2 with the stock configuration, 6.2x against what make currently produces.

The build flag change, which is separable

src/linux/Makefile currently appends -Os and links -fsanitize=address into every build unconditionally. This PR switches to -O2 (measured above: 1.24s vs 0.58s, a 2.1x difference on its own) and makes ASan opt in with make asan=1.

I have kept ASan easy to reach because it is clearly there on purpose, but having it in the default build also means make currently fails outright on any machine without a matching libasan, which is how I ran into it. Happy to drop this hunk if you would rather keep the current defaults; it is independent of the two interpreter changes and can be split into its own commit or dropped entirely.

Second commit

Build libzforth as a shared library, with and without tracing adds an opt in make libs target. make with no arguments still builds only the interpreter, so nothing changes for existing users. It is here because it carries one finding that is really a performance fix: zf_push() and friends are public API, so under -fPIC the compiler must assume the dynamic linker could interpose them and emits 70 PLT calls into the inner loop instead of inlining, costing the shared library a factor of two against the same code linked statically. -fno-semantic-interposition fixes it.

This commit is entirely optional and I am happy to drop it if a shared library target is not something you want in the tree. The first commit stands alone.

Verification

  • Byte identical output against current master for core.zf, mandel.zf, dict.zf, memaccess.zf and misc.zf.
  • Byte identical -t trace output.
  • Builds clean under the existing -Wall -Wextra -Werror -pedantic, with ZF_ENABLE_TRACE both 1 and 0.

James Hickman and others added 2 commits August 28, 2026 09:58
Profiling the inner loop with perf shows it is bound purely on instruction
count, not on stalls: a 65M dispatch benchmark runs at IPC 3.0 with 13k branch
misses out of 2.0e9 branches and 25k L1 misses out of 2.5e9 loads. The jump
table the compiler generates for the do_prim() switch does not show up in the
profile at all. The time goes into work done per instruction, so this removes
two pieces of it.

Decode single byte opcodes inline in run(). Most dictionary cells encode as one
byte, but every instruction fetch paid for an out of line call into
dict_get_cell_typed(), which was 22% of the profile on its own.

Only walk the return stack for the trace indent when tracing is actually on.
The loop over RSP(ctx) ran once per instruction regardless, and each iteration
reloaded RSP through ctx->uservar, which aliases ctx->dict and so cannot be
kept in a register.

Together these cut the benchmark from 7.44e9 to 5.22e9 instructions, 114 to 80
per dispatch, and the run time from 0.605s to 0.385s at -O2. Output is
unchanged on core.zf, mandel.zf, dict.zf, memaccess.zf and misc.zf, and the -t
trace output is byte identical.

Build the interpreter at -O2 rather than -Os, worth a further 2.2x, and make
the AddressSanitizer build opt-in with 'make asan=1' instead of linking it into
every build by default.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H676rBN7QwnkCHZgBdNKk8
Adds a 'make libs' target building two shared library variants from the same
source and the same zfconf.h. libzforth has the tracing code compiled in,
libzforth-notrace does not.

Compiling tracing in costs about 45% of run time even when tracing is switched
off at run time, because every trace site still tests the flag, and reading it
goes through ctx->uservar, which aliases ctx->dict and so cannot be kept in a
register. On the 65M dispatch benchmark the traced library runs in 0.36s and the
untraced one in 0.21s. The untraced library is also 12 kB smaller and does not
reference zf_host_trace() at all, so a program embedding it does not have to
define that callback.

The two carry different sonames and can be installed side by side. They are
interchangeable at the API and ABI level, since ZF_ENABLE_TRACE does not appear
in struct zf_ctx.

Both are built with -fno-semantic-interposition. zf_push() and friends are
public API, so under -fPIC the compiler otherwise assumes the dynamic linker
could interpose them and emits 70 PLT calls into the inner loop rather than
inlining. That alone was costing the library a factor of two against the same
code linked statically: 0.45s versus 0.22s.

ZF_ENABLE_TRACE in src/linux/zfconf.h is now only defined if not already set, so
the two variants can share one configuration header rather than duplicating a
file whose dictionary and stack sizes are baked into struct zf_ctx.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H676rBN7QwnkCHZgBdNKk8
@jameshickman jameshickman mentioned this pull request Aug 28, 2026
@zevv

zevv commented Aug 29, 2026

Copy link
Copy Markdown
Owner

Thank you, but I won't merge this though: zForth basically is only about the implementation .c file only, all of the other infra in the repo (makefiles, targets, etc) are just for demo purposes. No one is supposed to run or build zforth like this, asan is just there by default to make sure I don't make stupid mistakes, tracing is just there for debugging. The goal of zforth is to pick up the portable .c file and integrate it into whatever build system you already use for your project.

The same with the RPM package PR from earlier this week: I appreciate the effort, but this is just totally outside the scope of what zforth tries to offer. Ideally this repo would not even contain any build infra at all, no makefiles, etc, but that makes it hard for people to quickly give it a spin and see what it does.

I do see how the 1-byte fast path helps your case, but I'm not happy that it now bypasses the range check and duplicates knowledge that only is supposed to live in dict_cell_get_typed(). Also, personally I prefer simplicity, small size and correctness over performance.

However, if you happen to build and maintain a fork that has all of this, I'd be happy to link it from the README.

@jameshickman

jameshickman commented Aug 29, 2026 via email

Copy link
Copy Markdown
Author

@jameshickman

jameshickman commented Aug 29, 2026 via email

Copy link
Copy Markdown
Author

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants