Leveraging a blazing-fast runtime: a (new) Go backend for PureScript

News #19:

The project is now officially entering its cleanup and maintainability phase.

However, as I started diving deep into the compiler’s performance and architecture, I’ve been maturing a rather radical choice: rewriting the compiler backend in Rust (a temporary repo will evolve here). I’ve managed to completely avoid RAM overloads using a few tricks (.purmeta, etc.), but the further I go, the more I realize that there are things more important than that missing (see the list below).

Here is why I’m considering this pivot:

  • True parallelism: I want to fully leverage multi-core processing to parse and transform TAST files, etc. While we could theoretically bootstrap gopurs to compile itself and use Go’s concurrency, debugging a self-hosted AOT compiler that relies on its own generated code is a chicken-and-egg challenge I don’t really have the time to deal with, these days. Using a dedicated systems language like Rust gives us a direct and fearless parallelism out of the box.
  • Raw speed: I want the compilation times to be as fast as possible, even during single-threaded execution. E.g., Rust’s performance profile and insanely fast JSON deserialization are perfectly suited for crunching the tcorefn payloads.
  • Memory reliability: TAST never erases types, the AST graph we hold is really big. In GC-based environments, allocating millions of tiny AST nodes puts great pressure on the Garbage Collector. Rust allows us to bypass this entirely (e.g., using Arena Allocators), giving us a perfectly predictable, highly reliable, and performant memory model.
  • Natural translation: you might wonder if porting from PureScript to Rust means losing our expressive FP abstractions. Thankfully, the core of this compiler isn’t heavily reliant on advanced type-level programming. It is essentially a pure pipeline transforming massive ADTs. This maps flawlessly to Rust: it won’t require translating into complex traits or fighting the lack of HKTs. It just means writing a lot of enums and exhaustive match statements. I’m oversimplifying a bit, but it guarantees a smooth architectural transition.

I want to be transparent: this decision might delay the official release by at least a month. But I believe the long-term benefits for our ecosystem are well worth the wait.

Perhaps purust will reach a level of maturity where we can leverage it directly one day to write compilers with native parallelism. But as it stands, writing the tooling directly in Rust seems to be the most reasonable path to ensure the compiler is as performant as possible.

Note: a Purescript developer could certainly use Node.js for local development (no AOT compiling penalty), and Go or Rust for production, but my initial goal was a bit more ambitious than that: I’d like to enable a Purescript developer (who doesn’t work on frontend projects) to completely forget about Node.js and npm. Not out of any particular animosity toward this technology, but for the sake of convenience (only one runtime target, no prod surprise, etc.).

Note #2: purbo has been created, and it is a direct Rust port of @natefaubion’s brilliant purescript-backend-optimizer. All credit goes to him for the semantic architecture, this will just be a translation to leverage Rust’s parallelism and memory management. That is something gopurs will rely on, like other backends in the future.


Edit: A slightly more nuanced compromise was made. More details here.

News #20:

Oh… I was starting to doubt it, but in the end, the Wrapper/Worker pattern paid off in gopurs: we’re now down to 42ms… It really is a very strange feeling, like a never-ending logarithmic slide. But that’s a good thing. Benchmark updated here.

News #21:

I am considering that the best approach moving forward might be to have a coexistence of PureScript and Rust code within gopurs.

The idea would be to use Rust to compile the official, highly-performant release binary, while keeping a PureScript version to compile a dev binary and easily understand what the compiler does. This could ensure that the community can easily experiment, create proof of concepts, and submit PRs using PureScript. In this context, more time to compile is absolutely not a problem.

Nowadays, AI makes it more easy to “transpile” diffs if a contributor wants to patch the Rust mirror, without learning it in depth. (Of course, I could handle the Rust porting myself. Or purust could do this automatically in the future.)

Ultimately, gopurs is a tool for the PureScript community, and my goal is for it to remain maintainable by the community, not just by me. Rust would be there strictly for performance reasons (serde for JSON decoding, parallelism…), not to serve as the documentation medium for compilation methods. PureScript remains much more refined, concise, and its expressivity (which is one of its greatest qualities) makes it the perfect language for documenting and designing the compiler’s logic.

Also: currently testing the Scrutinee Fusion method (inspired by id3as and their purescript-backend-erl compiler).

News #22:

Latest work completed on Records, making the most of the power of Go struct-ures.

Deep Record Updates: 150 μs → 5 μs.

News #23:

A final optimization push has been made.

The gopurs-compiled code (~ 40.5 ms), which doesn’t hesitate to use every means at its disposal (e.g., deep monomorphization, stack-allocated structs, goto…), finally reached the symbolic milestone of being twice as fast as the Arista-compiled version (~ 81 ms), on single-thread tasks.

Just in case, I’m still going to try using gopurs.js soon to compile gopurs.go so I can use it as the official binary (and thus take advantage of its parallelism). Of course, Rust is cool and useful, but I’m still a little unsure about whether it really makes sense to use it right now. The debate is still in my mind…

News #24:

Another major milestone achieved: the total execution time on the core stress tests just dropped from ~40.5 ms to ~28 ms!

This massive 30% gain comes from a single, targeted compiler pass I just implemented, which I’d call “Thunk Fusion” (added a dedicated test for it in the official passing/ folder).

By deeply analyzing the TAST, gopurs can now detect when a Lazy thunk encapsulates pure, total integer arithmetic and is immediately forced by its consumer. When this precise pattern is matched, the compiler safely bypasses the dynamic closure allocation entirely and generates a strict, unboxed, tight loop in Go. It essentially synthesizes the “cheatcode” performance automatically.

Crucially, this does not break the Lazy contract or semantics:

  • It only applies to total operations (no side-effects, no risk of evaluating diverging effects prematurely).
  • It strictly requires an immediate consumer. If the Lazy value is passed around or stored, it gracefully falls back to a standard deferred closure.
  • Recursive LetRec blocks are explicitly excluded to maintain safe initialization semantics.
  • etc.

This perfectly embodies the AOT philosophy I was aiming for with this backend: letting the compiler cancel out the overhead of high-level functional abstractions (like abusing Lazy in a hot path) without forcing the developer to manually drop down to FFI or rewrite their code imperatively.

In doing so, the compiler becomes more and more forgiving of upfront design mistakes. A PureScript developer can write naive, unoptimized code (abusing closures or deferring execution unnecessarily) and the compiler will silently catch and clean up those architectural flaws behind the scenes, ensuring top-tier performance anyway.

I’m remaining cautious, of course. With optimizations this aggressive, there is always a risk of hitting unexpected edge cases in the wild. But the strict semantic guards I’ve put in place, along with the test suite passing flawlessly, give me more confidence.

We are now incredibly close to imperative and native Go (~27 ms).

Today is a good day. :sunflower:

News #25:

Once we crossed that threshold, it didn’t take long to achieve performance levels (25ms) that got better than native performance, for the first time! At this point, we’re not even that far off from native Scheme code (20ms) or native Rust code (16ms).

That’s the magic of compiled code: it doesn’t care about being as readable as hand-optimized code. It can do whatever it takes to run as fast as possible.

What’s great is that any big compilation achievement unlocks new ground for optimizations that were already primed to act, on code still trapped in a shell waiting to be dissolved.

1 Like

Whoa, pretty soon you’ll be whizzing past zero into ahead-of-time execution. Very exciting stuff, hats off to you!

1 Like

Thank you, Erik. That’s a beautiful message that I’ll hold close to my heart when other challenges arise.

1 Like

News #26:

Quick note to say that I had to run through all the test suites (modules, official ones, etc.) again after the recent performance optimizations. A lot of things were broken. Everything is green again.

The cleanup is still underway.

1 Like

News #25:

@erikagain I didn’t think we could get this close, but you might be right!

gopurs is now down to about 13ms for the full benchmark suite, compared with roughly 11ms for the hand-written Go implementations (which I also optimized: 27ms -> 11ms). There are reasons to believe that the compiled code will perform on par with hand-written native code, once again.

The 10ms mark is starting to look within reach. :desert_island:

News #26:

Okay, just like with purust, I’m now struggling to gain microseconds. It’s no longer in the millisecond range. So this is no longer a priority.

All my efforts are once again focused on the official release of gopurs. I can now see the light at the end of the tunnel.

1 Like

News #27:

Good news! :blossom:

I’ve decided to keep the backend written in PureScript and compile it to Go using gopurs itself, rather than rewrite it in Rust. That will also be the highway to maintainability for PureScript devs.

In fact, I’ve underestimated the progress made so far: gopurs can actually compile itself now, and the compiled code is super optimized.

It’s no longer just a promise: it’s already working. I’m compiling the small real-world project I usually talk about, in Go, using a Go binary from gopurs (instead of the usual JS one). We can now leverage its parallelism, speed, etc.

Using gopurs to build itself also gives us a substantial real-world workload. Improvements to code generation, the runtime, and the Go FFI can then benefit both the compiler and other PureScript applications.

Cheers! :heart:

1 Like

News #28:

Compilation time has been reduced by a factor of 5, and continues to improve. On a project like this one (always the same example), it used to take me 20 minutes; now it takes 4 minutes. (I’m referring to the first pass, not the incremental one.)

Another very positive sign: the official tests, as well as all the tests for each module, run successfully when gopurs is compiled in Go rather than in JS.

1 Like

News #29:

Go is now faster than JavaScript in terms of compilation time on several modules. Dogfooding (i.e., self-compilation) therefore seems to be paying off. Sometimes, the ideal of “10x faster” (like TypeScript v6 → v7) is reached (e.g., gopurs-functions), sometimes not (for now). Other gopurs-* modules will be added very soon. And perf improvements are still ongoing. (Note: we will be able to add public projects by developers other than myself, to this table.)

For b8x, 4 min → 1 min. See there for details.

Anyway… harder, better, faster, stronger. :upside_down_face:

Added a special benchmark for compilation here.

News #30:

The Rust-hosted version of gopurs, built using purust, now becomes #1 (comptime perf).

Across b8x and 50 library benchmarks, the sum of backend compilation medians is 18.8% lower than with the Go-hosted version, with byte-for-byte identical generated Go output.

This makes Rust a strong candidate to become the default host for all backend executables: gopurs, phpurs, javapurs, sharpurs (not just purust).

The idea is to compile each backend’s PureScript implementation into a native executable via purust, while keeping its own output target. E.g., a Rust-hosted phpurs would still generate PHP.

It’s still dogfooding, just on a broader scale: using one PureScript backend to build the others.

Native builds of the remaining backends and cross-platform distribution are the next steps.

1 Like

Any interest in a Zig backend (or a fork/optimization of the Cpp one if you’re up for it)? Not having linear types in Purescript makes the Purescript to Rust abstractions an awkward mismatch, IMO. I’d be interested to see an app that uses it to its full potential though and acknowledge I could be wrong.

Thanks for the Zig/C++ suggestion. I had previously tucked that away in a remote corner of my mind, because I’m primarily focused on making popular/trendy ecosystems accessible to PureScript developers, and (vice-versa) making PureScript accessible to developers coming from popular ecosystems. My initial view, given the sheer amount of work that still needed to be done, was a bit oversimplified at the time: “Zig is still immature in terms of its ecosystem, and C++ is obviously mature but a tiny bit of a hassle to manage when it comes to packaging/FFI”. Your suggestion makes me want to revisit those possibilities in the context of dogfooding, where faster compilation would benefit users of the existing backends. I’ll try a few specific experiments. Maybe I’ll even share some findings soon.

I see why the absence of linear types in PureScript raises questions about mapping its abstractions to Rust. It’s a topic that has always interested me, and I’d like to see similar capabilities explored in PureScript someday. I really like Haskell’s approach to linear types. Still open questions, in my mind.

For now, I’m working on that mapping in the backend and runtime. In purust , sharing is handled explicitly through reference counting, with atomic reference counting in threaded mode and copy-on-write where needed to preserve PureScript’s immutable semantics.

The backend also tracks which values are still needed, allowing it to emit moves instead of clones where possible, borrow in supported cases, and reuse certain allocations when a runtime uniqueness check succeeds (that check matters: a variable’s last use doesn’t necessarily mean there are no other aliases). This approach doesn’t require linear types in PureScript, although reference counting and allocations still have costs to measure and optimise.

So purust currently makes these decisions through its own analysis of the transformed backend IR, and runtime checks (that’s the cost).

I’ve also started enriching the TAST with optional usage annotations, aiming to make backend analyses more straightforward over time: bounds on direct uses of local bindings, potentially escaping use contexts, and last-local-use markers. These don’t establish exclusive ownership or absence of aliases (yet). Still an experiment at this stage.

For a substantial workload, gopurs itself is one concrete example: its PureScript implementation now runs as a Rust executable built with purust, compiling b8x and the library projects mentioned above to Go. That gives me something useful to evaluate beyond small benchmarks. But yes, we need more!

Developers will either be able to provide me with some PureScript code snippets that I’ll be able to test on my machine, or will be able to test them themselves after the nixification (one of my priorities today).

At this stage, purust remains experimental TBH, and using it to host the other backends is still a candidate approach. The results are quite encouraging, but I haven’t settled on it as the definitive solution. There are still open questions, and I want practical evidence across different workloads to guide that choice. In any case, it’s a great excuse to improve its performance, which will in any event benefit those who use it as a backend and a tool for leveraging crates (typically with a PureScript core).

To be continued…! :mag:

Supplementary reply:

I plan to make Rust one of the supported backends in b8x, a project developed for an actual client. It is already done, in fact. But I need to improve one or two things before making it public.

Definitely awkward.

And that’s exactly why I saw great potential in it. Let me answer that precise point, on a less technical level, but more exhaustively (because it is quite counterintuitive, and important):

When I was a kid, I liked opening (real) locks with pieces of wire. (I’m not talking about the time I found myself in deep shit and got through it by using a storage cupboard key to sleep in neighbouring apartments that had stood empty for years.) That’s precisely the lesson I take from those experiences: sometimes, the unimaginable is much closer than we think. Much much closer. The single blocking resource is… taking time to try it.

A key’s profile can give you access to your neighbour’s apartment simply because your landlord had the locks supplied by the same company: the key fits the same opening, and all it takes is forcing the turn a little and exploiting a known issue with mechanical play, so that some of the internal cylinders stay unlocked while you unlock the others by pulling and pushing. The details of these mechanical weaknesses are a bit more known nowadays (yet, in Paris for example where I live, 70% of locks are still insecure…).

Another example… Sometimes, you’ll find a reinforced door with a camera that activates automatically when the outside handle is used, but not when the inside handle is used: then all it takes is passing a hooked wire underneath the door and pulling… Three dollars can defeat a system costing several thousand.

In fact, the single subject of locks is in itself full of examples of the kind.

Houdini taught us a lot. :tophat:

I’m an orphan, and I went through horrific things before finding my way back into society. Those circumstances pushed me into doing things I would never otherwise have done, or even tried to do. That’s what I learned from the streets and from urban exploration (a.k.a. urbex): discovering something somewhere because nobody goes there: too creepy, dangerous, dirty, etc. The human world is, above all, a world of psychological forces. (I’ve found some unbelievable things by climbing over the walls of abandoned properties…). These are real forces. They’re not just hot air. But they’re psychological in nature, and they are not the final word.

I didn’t develop a taste for stealing, and I don’t encourage anyone to steal anything. There are many restrictions, and they exist for good reasons. But I certainly developed a taste for shifts, for surprises, for changes of perspective: those moments when something becomes possible because the way the problem was framed was the real problem.

Of course, it doesn’t work every time. It’s simply a way of approaching life. When you said “awkward”, my first reaction could have been to take it personally and (wrongly) interpret it as an expression of passive aggression. But I don’t see it that way, because that’s precisely the starting point. And you’re right to say the words. It’s strange, so it’s interesting to tinker with, to poke and prod. That’s exactly what inspired me to create phpurs and gopurs, too. I’m actually more interested in things that seem strange at first glance.

Because if it doesn’t work, we gain a better understanding of the lock’s inner workings. If it does work, it isn’t just a lock that opens: it’s a whole world.

So… thank you for bringing that up. It gives me a chance to introduce myself a little, on a personal level. I hope I’ve provided a better glimpse into my journey and the meaning behind every step I’ve taken in these projects, that I’m modestly trying to see through to completion. As you can see, intuition remains the driving force.

For now, I’m still making do with whatever I have on hand: the conviction that PureScript has made things much easier with its strictness, that its level of generality and abstraction allows us to do a lot of things at every level (including addressing the specific nuances of each language), and so on.

In a way, I actually think I’m more American than French when I act that way. (In France, Americans have a reputation for having one of these qualities: empiricism, concrete answers first).

We’ll see where this takes us.

1 Like

News #31:

purust is starting to come close to the ideal goal of x0.1 (10x faster), for the gopurs binary. I’ll pause the compiling perf improvements for now.

I’m currently working on improving FFI usability. My goal is to let developers use native types and ordinary function signatures on the Go side, while the backend handles conversions, currying on the PureScript side, and the necessary runtime wrappers. I had already started this step quite a while ago; I just hadn’t finished it.


@harryprayiv For purust, that’s the same: the goal is to keep the handwritten FFI idiomatic, with the backend handling as much of the adaptation as possible. I’m exploring how generated FFI adapters could enforce some ownership constraints at runtime, such as preventing a resource from being consumed twice, while keeping the handwritten Rust FFI idiomatic. Recovering comparable compile-time guarantees would require more work, though (but I’ll investigate).

Edit: The discussion has been extended here: Leveraging modern low-level: a Rust backend for PureScript - #29 by 0x000000000000000000