Leveraging modern low-level: a Rust backend for PureScript

Hello everyone,

As some of you might know, I recently worked on phpurs and gopurs. Building these backends has been an incredible journey, but it also highlighted some hard boundaries. Specifically, gopurs showed me the limits of Go when it comes to implementing advanced memory management techniques like Perceus and FBIP (Functional But In-Place), inspired by Koka/Lean. The gopurs project has achieved incredible performance, exceeding my expectations (as explained at length in its thread), and that’s a huge victory, but I think we can still do better.

I am currently in a sabbatical phase of my career, without an official position, and I am ready to invest the time to create pure tools that will sustainably serve the ecosystem. I absolutely love PureScript, and I believe it deserves the best tools.

Looking at the current landscape, the planets are perfectly aligned to attempt something experimental but ambitious. A native Rust backend: purust.

Here is why I think the timing is right:

  • PureScript is strict: Unlike Haskell, PureScript’s strict evaluation makes it the perfect candidate for Perceus-style static Liveness Analysis.
  • The backend-optimizer is proven: Having a flattened, uncurried, and optimized AST ready to be consumed removes the heaviest burden of writing a backend.
  • TASTs are available: Having the type information attached to the CoreFn is the missing piece of the puzzle. Still local (on my computer), but completely unlocked the perfs on the Go compiler.
  • The Value architecture worked: gopurs proved that dynamically managing Boxed vs. Unboxed data structures works well. In Rust, we can take this even further.
  • AI accelerates the process: Today’s AI tools allow us to iterate, prototype, and test complex compiler ideas much faster than before.

At first glance, Rust might seem like a hostile target. The Borrow Checker, lifetimes, and the heaviness of standard Rc or Arc seem to clash with massive graph sharing and closures. However, the initial roadblocks for phpurs and gopurs also seemed insurmountable at first, yet they were overcome. By generating highly specific unsafe Rust code encapsulating a custom PerceusBox, we can bypass the Borrow Checker entirely and leverage LLVM’s raw power.

I love mathematics, and in math, we know that when abstract shapes seem to fit together, even from a distance, intuition is often the only argument that leads us to incredible new territories. That’s how I feel today, thinking about this potential PureScript-to-Rust compiler.

Engineering often begins with a dream, much like reading Jules Verne before building the submarine.

I am currently working on the foundations of this project. The repository will be made public soon, provided that the initial proofs of concept, facts, and benchmarks prove the architecture is credible.

It is absolutely worth a try. In the worst-case scenario, I will potentially save time for anyone else who might have the same idea, by explaining and documenting the limits of this exercise.

At best, a whole new avenue of joy opens up for PureScript.

1 Like

Hi. Wow. I can hardly believe my eyes. You’re building yet another backend.

The memory-ownership argument makes sense to me: a tracing GC does rule out precise Perceus-style ownership, and that’s a real limit rather than a design mistake on your part.

What I don’t follow is why that points at Rust. If the plan is unsafe code with a custom PerceusBox bypassing the borrow checker, then lifetimes, ownership and the aliasing discipline are all switched off, and what’s left is LLVM with Rust syntax on top. Koka, which is the reference implementation of the technique you’re after, targets C. purec already exists here and could use some love perhaps. So what does Rust give you once the checker is out of the picture that C or Zig or direct LLVM IR wouldn’t?

The part that worries me is that Rust’s unsafe is stricter than C rather than looser. &mut emits noalias, and provenance and Stacked Borrows make aliasing patterns that are fine in C into UB that LLVM is entitled to miscompile. A refcounted aliased object graph is close to the worst case for that. If you go this route I’d want to know the soundness plan early: is the whole thing running under Miri in CI, and what are the documented invariants on PerceusBox? That’s the sort of thing that’s very hard to retrofit.

Separately, and this is the thing I’d most like to see regardless of which backend wins: tcorefn is still local on your machine and it’s now a dependency of three projects. It’s the most reusable thing you’ve built and the only piece that survives a pivot. Even a rough patch and a format sketch in a public repo would let other people build on it and would make both backends reproducible.

And a straight question, no judgement in it: is the gopurs release still happening? You mentioned cleanup, Aff and an official release yesterday, and I’d like to know whether to treat gopurs as something to plan around or as a finished experiment.

2 Likes

Hi @harryprayiv,

Thank you so much for this high-quality feedback. These are exactly the kind of technical challenges and hard questions I was hoping to discuss. You hit the nail on the head on every point. Again, I think we’re kind of synchronized on the questions to deal with.

Let me start with your last question because it sets the context for everything else: gopurs is my absolute priority.
The official release, the cleanup, and the Aff integration for gopurs are happening. You can treat it as a project to plan around. purust is currently an exploratory side-project (a non-priority baby that I tinker with), meant to push boundaries and see if the Perceus architecture can theoretically land in the PureScript ecosystem.

Regarding your technical points, here is my thought process:

1. Why Rust over C, Zig, or direct LLVM IR?
You are completely right: bypassing the borrow checker leaves us with LLVM + Rust syntax. If memory ownership was the only metric, C or Zig would give me much more freedom.
However, this choice is entirely driven by the Ecosystem and Developer Experience (DX).
By targeting Rust, purust gets to leverage Cargo for seamless dependency management and cross-compilation. The libs are really awesome and numerous. I’d like to open a door to this growing world in the most direct and seamless manner. It gives us a golden path to map PureScript’s Aff directly onto the Rust async ecosystem (like tokio), and allows developers to easily write FFI bindings to world-class crates (serde, axum, etc.). Rust is treated here as a high-level, portable assembler equipped with the best (low-level) package manager in the industry (IMHO).

2. The Strictness of Rust’s unsafe and Aliasing (UB)
This is the most critical hurdle, and you are 100% right to point out that Rust’s &mut emitting noalias is stricter than C. A refcounted aliased object graph is indeed a minefield for Stacked Borrows.
The soundness plan relies strictly on the mathematical guarantee of Perceus/FBIP: we never form a Rust &mut reference unless the reference counter is exactly 1 (meaning we have absolute proof of unique ownership). For any shared state (ref count > 1), the data is accessed via raw pointers (*const T) without ever upgrading them to Rust references.
But you are entirely correct: human intuition is not enough here. Running the generated code under Miri in CI will be a mandatory requirement to mathematically prove the soundness of the PerceusBox invariants and avoid LLVM miscompilations. Still playing with it.

3. The tcorefn (TAST) format
I couldn’t agree more. The typed AST is indeed the most universal and reusable piece of architecture that came out of this journey. It completely unlocked the performance of the Go compiler for me. Releasing a sketch/draft of this format in a public repo so that other people can build on it is definitely on my roadmap. It survives any backend pivot and should belong to the community.

Thanks again for the peer review and the Miri advice. It’s incredibly valuable to have this kind of stress-test before diving too deep into the code! You’re definitely a friendly ally to me, in these experiments.

2 Likes

And thank you for your encouragement and compliments. It’s the kind of kindness that I find all too rare in some corners of our field, and that I truly appreciate. It only confirms the positive spirit that exists in this PureScript community.

1 Like

No problem. Glad to help.

2 Likes

News #1:

It’s not a victory yet, but we can still call it progress, since the backend is now able to run all tests in this benchmark. I was at about 10 seconds to start with. The usual path of successive optimizations is proceeding as planned.

1 Like

News #2:

Just fixed a huge memory bottleneck. Total core benchmark execution dropped from ~2.4s to ~484ms.

By leveraging the TAST to avoid unnecessary deep clones on ADT field accessors, native memory management for functional data structures gives:

  • Red-Black Tree: ~1.98s āž” ~61 ms (32x faster)
  • List Processing: ~1.2ms āž” ~82 μs (15x faster)
  • Prime Sieve: ~8.2ms āž” ~615 μs (13x faster)

I think purust can now be made public.

Still a lot to optimize.

1 Like

News #3:

purust is now way closer to the official JS backend: 484ms → 190ms (vs JS: 125ms)

Still a lot to do. But that’s getting cool.

News #4:

All right! We’ve fallen below the official JS backend threshold (i.e., 125ms), with a total of 119ms.

To be continued :green_apple:

News #5:

Arista :white_check_mark: (81ms)
Scheme :white_check_mark: (46ms)

Benchmark here.

Of course, the final major step of cleaning up what was co-produced with AI still remains. But since I’ve already started that process with gopurs, it should go smoothly. For now, I’m continuing to dig deeper and deeper, to see how far we can go (which is much faster now that almost all optimization techniques have been tested with gopurs: I can directly reproduce them).

News #6:

Finally happened: purust (21ms) is now ahead of gopurs (24ms)…

The Koka-n way of life is giving good results.

I haven’t mentioned this backend yet, because I don’t want to take over the forum, but javapurs is the new champion of the benchmark (for now): the new goal is to reach its metrics. Happily, there’s still things to do… (e.g., finish the Perceus plan, which is still only modestly addressed).

An exhaustive (and a bit dirty) progress table is available here. The goal is to get white/green everywhere in the first 5 columns. The first lines are prioritized.

News #7:

Done: purust is now in first place (17ms) :white_check_mark:

News #8:

I’m not going to ask an AI to check my English in this post. Not this time. Pure joy takes over.

For the first time, a backend is close to reach the highly symbolic 10-millisecond barrier. It might not quite make it, but the mere fact that it’s so close is amazing in itself: purust (12ms) is now ā€œfarā€ ahead of gopurs (24ms, twice as fast!) and javapurs (17ms). The closer you get to 0, the more each millisecond feels as a victory.

Is this a ā€œEureka!ā€ moment? Perhaps something akin to an Act II for PureScript? That’s how I feel about it, personally.

The benchmarks aren’t perfect, nor are the numbers, nor is the code. In fact, nothing really is. The lab is a mess, and the to-do list is a mile long. I’m aware of this, and I thank everyone following these developments for their patience and tolerance in the face of the imperfections left by my AI-copiloted experiments. I’m aware they exist. Above all, I’ll remember your support and enthusiasm.

I’m trying to maximize the reach of the momentum and clear the ground as best I can before tackling precision engineering. Even though I don’t currently hold any official position, my time and energy remain fairly limited, as I have other important commitments on the side (volunteering, moving, getting started with job seeking, etc.).

But the essentials are in place.

The chemical reaction finally seems to be working without altering the two compounds: Rust’s speed is preserved, as is PureScript’s concise and pure expressiveness.

And of course, the doors of crates swing wide open: leverage a low-level ecosystem that is both rich and easily accessible.

Next steps: official tests, module-by-module tests (e.g., Aff), real-world conditions, etc.

Edit #1: pscpp seems to take 9.5ms, that’s my next goal.
Edit #2: #1 turned out to be false

1 Like

News #9:

The heavy lifting for Aff support in purust is done, and it works wonderfully. I’m really happy with the result. Tests 100% green (e.g. purust-aff) :green_circle:

Under the hood, purust-aff is natively built on top of tokio (using its multi-thread capabilities). That’s rather similar to what’s been done with goroutines for gopurs.

This gives us full native support for the entire async API: Fork, ParApply, Sequential, Fiber cancellation, Bracket for safe resource management, etc. (lots of things).

There are still tiny details to fix, but this will be pushed soon. The surrounding ecosystem (like purust-js-promise and purust-js-promise-aff) still needs to be ported. And as usual, a cleanup will need to be done at the end.

I’ll test that soon, via a benchmark which is mainly dedicated to concurrent/parallel metrics.

Also: I’m currently trying to get tests green on a real-world project (website here). I’ll see how it handles the FFI and the Aff spec test suites.

News #10:

Done. The (unit & integration: psql, rmq…) tests of a real-world project are now 100% green with Rust. :green_circle:

News #11:

Added to the concurrent/parallel benchmark.

There is the same factor of 2 between Rust and Go, and overall, Rust is about 20 times faster than JavaScript.

News #12:

And there you have it: Rust has finally caught up with C++: under the 10 ms barrier!

… or so I thought.

It turns out that was a phantom goal. Fatigue and all those late nights grinding on these backends led me to make a pretty crazy mistake with the units of measurement. In reality, the pscpp (C++) metrics are actually closer to a full second, making it about a factor of ā€œonlyā€ 2x faster than Go (psgo). Given the fact that both psgo and pscpp come from the same work (purescript-native), this actually makes a lot more sense and is much more consistent with my empirical observations so far, which also show roughly a factor of 2 difference between Rust (purust) and Go (gopurs). (This also aligns with the initial intuition that TASTs are of great value because they provide all the necessary type information to AOT backends: pscpp is based on untyped ASTs.)

In short, my unit mistake at least motivated me to push hard. The good news is that it leaves us with a very positive takeaway for purust: it’s incredibly cool to have such a high-performance low-level backend. 100x faster than pscpp :slight_smile:

That being said, I’m going to take a break for a few days to catch up on some much-needed sleep and recharge my batteries.

CheerzZz… :zzz:

1 Like

News #13:

Just out of curiosity, I added Koka to the benchmark suite to see where our recent advancements stand against a language now famous for its raw functional performance (thanks to its FBIP and Perceus architecture).

The results are very positive: purust is 3x faster. Koka achieves a better score than pscm (Scheme), which is impressive for pure functional code. It really puts into perspective how effective the AOT/TAST approach is when translating high-level semantics down to flat imperative structures.

Eventually, I’ll add more stress tests to confirm this initial concrete feedback on the performance comparison with Koka.

I wrote the native implementation with the help of an AI, as I am not at all familiar with the language (though I find it fascinating, but that’s another story). If anyone here knows Koka, feel free to review the benchmark script to check if it’s idiomatic or if it can be improved. You can find the code here.

News #14:

Added Haskell, OCaml and C to the benchmark. The reference is now ~ 8.7ms, and purust has (almost) reached it. (Microsecond-level optimizations still need to be done.)

New stress tests are coming (n-queens, breadth-first search…)

News #15:

Last, but not least.

Fable has been added to the benchmark: purust is 10x faster. (Further evidence that the major phase of optimizations is likely behind me.)

(For those who can’t zoom in: purust is < 10 ms, while Fable takes more than 100 ms.)

1 Like