Less is more

Real Floats Using OpenVM's Custom Extensions

Last post talks about my failure optimizing softfloat. As we discussed, fixed-point math is not a universally agreed solution. Do we just give up and accept not having floating point numbers on chain?

I mentioned OpenVM at the end of the Jolt post. I mentioned that OpenVM on 32-bit is a blocker. Boy, was I wrong. There were two things I missed:

So I experimented with OpenVM, and ported the same ray tracer demo to OpenVM. First, here are some numbers:

| VM                                                    | Cycles (IPC = 1)   | Speedup |
|-------------------------------------------------------|--------------------|---------|
| Jolt (rv64im + softfloat)                             | 166,905,547 cycles | 1x      |
| Jolt (rv64im + softfloat + non-float optimizations)   | 122,043,178 cycles | 1.36x   |
| OpenVM (rv32im + softfloat)                           | 123,294,825 cycles | 1.35x   |
| OpenVM (rv32im + softfloat + non-float optimizations) | 99,083,769 cycles  | 1.68x   |
| OpenVM (rv32imf)                                      | 29,591,216 cycles  | 5.64x   |
| OpenVM (rv32im_zfinx)                                 | 29,467,504 cycles  | 5.66x   |
| OpenVM (rv32imf + non-float optimizations)            | 5,283,763 cycles   | 31.58x  |
| OpenVM (rv32im_zfinx + non-float optimizations)       | 5,147,995 cycles   | 32.42x  |

In a nutshell, by building a custom extension, we can have native floating point numbers (provided by RISC-V’s F extension or zfinx extension) in OpenVM. In the best case, we can achieve a ~32x speedup compared to softfloat solution! It’s worth mentioning that raw RISC-V instructions executed are not the full picture for performance. Proving floating point instructions might be much more expensive than proving integer instructions. But personally, I would still consider these usable results.

For comparison, here’s what I get on native x86_64 environment (I’m testing on Intel Core i5 11500):

$ make build-native USE_FLOAT=true
$ sudo perf stat ./build/raytracer_native > image.ppm
Done.

 Performance counter stats for './build/raytracer_native':

                 0      context-switches                 #      0.0 cs/sec  cs_per_second
                 0      cpu-migrations                   #      0.0 migrations/sec  migrations_per_second
               145      page-faults                      #  87126.4 faults/sec  page_faults_per_second
              1.66 msec task-clock                       #      0.2 CPUs  CPUs_utilized
            26,577      branch-misses                    #      1.6 %  branch_miss_rate
         1,667,836      branches                         #   1002.2 M/sec  branch_frequency
         6,134,926      cpu-cycles                       #      3.7 GHz  cycles_frequency
        10,074,298      instructions                     #      1.6 instructions  insn_per_cycle
                        TopdownL1                        #     37.0 %  tma_backend_bound
                                                         #     12.9 %  tma_bad_speculation
                                                         #     16.4 %  tma_frontend_bound
                                                         #     33.7 %  tma_retiring

       0.001997706 seconds time elapsed

       0.001034000 seconds user
       0.001036000 seconds sys

On native the program takes ~10 million instructions as well. It’s comparing apples vs oranges of course, but ~5 million RISC-V instructions are needed on RISC-V, maybe we have gotten rid of most overhead.

Note: for simplification, I’m only building single-precision floating numbers now. Earlier Jolt sample from previous post used double-precision floating numbers. I have since added single precision support in the Jolt demo as well. All numbers in this post are gathered running single precision support. So if you were puzzled that previous post shows ~180 million cycles for Jolt example, but here we show ~166 million cycles, it’s due to the precision difference.

OpenVM

OpenVM was developed primarily by Axiom. One of its key features is a modular no-CPU architecture: it does not have a clear definition of CPU. Even the core rv32im ISA is designed as an extension. All data flow through shared buses. This makes it quite easy to extend the ISA, adding new instructions. The extension building process is actually well designed, I didn’t run into any hassles during the whole experiment. Everything is quite neat.

I do have a question regarding proving side as well: will hosted proving infrastructure, like the Axiom Proving API, provide custom extension support as well? It will be really great if proper sandboxing can be designed so proving works with custom extensions as well. We will have to wait and see.

That said, for games, I think it might be in its particular sweet spot: chances are the gamer running games will already have a not-too-shabby GPU, so when proving is really needed, one can just use their own GPU to generate the proof. It might take longer on a single GPU but it might be fine as we are talking about a single proof. So maybe proving with custom extensions is not that critical.

openvm-floating

Following the official doc, I started building a custom extension. I must admit I’m more a VM person and know less about zero knowledge proof theory. So I relied on GLM 5.2 to help me code most of the extensions. Interestingly, while I was building the extension, Kimi 3 came out. I then had the two LLMs working on the extension in rotating shifts. The final result is here, both models reported that I didn’t get soundness fully right, but at least it can run enough code, and some of my proving tries do work now. So praise the agents :P I do want to get into ZK theories so I can maintain this library myself someday, but not today.

F vs zfinx extension

I actually started out with openvm-floating implementing RISC-V F extension, the single-precision floating point part. It turns out there is also zfinx extension, which is similar to F, except that zfinx (z-f-in-x) runs floating point operations in x registers. F extension runs in f registers, thus needing transfers between x and f registers, as well as transfers between f registers and memory.

They do not differ too much, so I asked agents to implement both. And the numbers showed that for ray tracer, zfinx actually has slightly better performance. I guess for the simpler ray tracer example, register pressure hasn’t become a problem, and the extra ABI requirements from F become visible overhead. In the future I will try more complicated programs, we shall see then which extension yields better performance.

Putting Together The Port

Now we can put together our port of ray tracer to OpenVM.

OpenVM’s RISC-V Architecture

OpenVM’s spec is really good (I wish I could write documentation of this quality back when I was working on CKB-VM -_-), skimming through it gives you almost everything you need to know about OpenVM’s RISC-V architecture. While I primarily build programs in C / C++ these days (Odin I’m coming), this Rust Frontend page is full of gold:

musl changes

On top of the earlier Jolt work, it didn’t take many changes to add OpenVM support in musl. But there are certainly some things I ran into:

Outputs

There is one thing I ran into: while Jolt offers a configurable memory region for bigger outputs, OpenVM limits the output values to a couple of u32 values, configured by num_public_values. I didn’t find any discussion about making this number really big, e.g, 4194304 so we can generate 16MB of output data. So I didn’t risk doing it.

Since I can tweak the heap configuration, I did the following setup instead:

+-----------------------------------+ 0x20000000 (512 MB)
| Reserved Memory for Output        |
| Ray Tracer Image (~32 MB)         |
+-----------------------------------+ 0x1E000000 (480 MB)
| ^                                 |
| | Heap Grows Upward               |
|                                   |
| Heap Region                       |
+-----------------------------------+ `_end`
| ELF Code / Data / BSS Segments    |
+-----------------------------------+ 0x00200800 (~2.002 MB)
| xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx |
+-----------------------------------+ 0x00200400 (~2.001 MB)
| Stack Region                      |
+-----------------------------------+ 0x00000000 (0 MB)

I remember first reading about this trick when I was diving into SP1 code. SP1’s guest program would reserve 16GB memory region at the very end of memory address space, but for storing read input values. As we have the need, there is nothing stopping us from leveraging the same trick for output values by OpenVM.

But I’m really new to OpenVM, I’m not 100% sure if this is free of quirks. Maybe there are some other restrictions I’m not aware of.

Running Everything

Now we can start running everything, the README file of the repo already contains the steps, so I won’t repeat them here.

There is only one gotcha I want to point out: for now openvm-floating only has single precision support, so double precision operations will still be implemented via softfloat. When you are running the demos here, USE_FLOAT=true is a mandatory flag.

Profiler

As of the time I tried, there is no way to expose pc traces from OpenVM. I had to add a patch, exposing a hook so I can peek into runtime environment. With it we can build a profiler supporting inferno based flamegraph, pprof and also samply. For example, here’s a pprof chart of the final optimized program, after optimizations from the next section are also applied:

Profiler Chart for OpenVM Guest Program

Optimizations

When softfloat no longer is a bottleneck, slow code from the integer portion of the ray tracer program will then be visible:

Looking at the number: using rv32im_zfinx before those optimizations, ray tracer takes ~29.5 million cycles to complete. When those optimizations are applied, ray tracer now only takes ~5.2 million cycles. Even when floating point operations become native, we can still obtain another 5.72x speedups by profiling the code. With softfloat dominating the stack, it’s hard to pinpoint those optimizations.

Picturing ZK Powered On-chain Games Together

Now we have usable native floating point numbers supporting on-chain games, I’m picturing the following setup:

  1. Native OpenVM floating point extension is implemented to use the exact same behavior as native x86_64 / aarch64 floating point operations. As JoltPhysics has demonstrated, this is actually doable.
  2. In the happy case, games run on native x86_64 / aarch64 code, using native floating point instructions for maximum performance.
  3. At certain intervals (e.g., 1 minute based on Teeworlds / OHOL convention), the server updates the current world state hash on-chain, and submits player inputs for current interval either also on-chain, or to a cheaper DA solution.
  4. All players can monitor on-chain values. When a mismatch is detected, a player can build a proof for the last chunk’s gameplay, and submit it on-chain. Since there’s no point running a rigged game, the player can leverage GPU to accelerate proving. The proof effectively validates the last chunk’s gameplay, proving misbehavior of the server. The game can then settle or revert and continue from a previous point, depending on the game design.

Determinism is crucial here: in the optimal case floating point operations run natively on CPU; but when proving happens they run on OpenVM with floating point extensions to build proof. Both should have the exact same bit-by-bit behavior.

With this combination of optimistic-based and ZK-based design, we can afford to build games at a much larger scale, but still be able to validate the full game on-chain.

OpenVM really shines here with the custom extensions. We can build a native floating point extension into it, without forking OpenVM at all. It certainly won’t be an easy task to build a fully secure floating point extension (I’m still on it), but the fact that this even works, surely looks like a miracle to me.

A Spectrum Of Solutions, For Games

For the past few months I’ve been experimenting with different setups. I started with Teeworlds and OHOL directly running on CKB-VM, to two ZK VMs running ray tracers with different floating point operation support. The reason for those efforts is that games vary a lot, different designs require different architecture and implementations. I want to have a spectrum of solutions, covering as many kinds of games as possible. If one game designer wants to build on-chain game with their own ideas, we should have a set of architectures suiting whatever design they want to go for.

While my earlier demos with Teeworlds and OHOL show how to build simpler 2D games on-chain, I think we can do more using OpenVM ZK VM with floating point extension. You might have already noticed that I talked a lot about JoltPhysics, this is actually my next target: with floating point extension powered OpenVM, I want to port JoltPhysics on chain and run it efficiently, and the journey does not stop there: as JoltPhysics is now the official physics engine of Godot, can we go one step further, and have a full Godot-powered game running on-chain via a ZK VM?

If you are interested, stay tuned. I will share what I find.