Kyo vs cats-effect: promising numbers and one big surprise
I find Kyo one of the most interesting projects in Scala right now. I get the feeling that some of the complexity we accept might not be necessary.
That feeling is easy to get excited about. I wanted to see some numbers too. This is the first post in a series about Kyo and cats-effect, starting with a small synthetic war.
I compared cats-effect 3.7.1 with Kyo 1.0.0-RC6, using fs2 3.13.0 on the cats-effect side for streaming.
Getting a fair comparison took more work than I expected. Two methods can look equivalent while doing different things with batches, buffers, or failures.
The suite passed 984 correctness checks. I ran JMH with three JVM forks and identical heap settings on one Mac with JDK 25.
Both sides use the same worker strategy and stream batch boundaries. The success-collection row uses parallel attempts and filtering, rather than native Async.gather. The stream rows do not directly compare parEvalMap with mapPar.
These are selected results with very little work per element. Ratios compare mean throughput; CE means cats-effect.
| Construction | Higher throughput |
|---|---|
| Uncontended permit acquisition/release | Kyo ~5.02× |
| Queue-backed stream batches | Kyo ~3.70× |
| Sequential fiber spawn/join | Kyo ~2.75× |
| Parallel stream batches | Kyo ~1.93× |
| Integer queue, one producer and consumer | Kyo ~1.45× |
| Bounded workers | Kyo ~1.13× |
| Parallel attempt and success collection | Confidence intervals overlap |
| Left-associated suspended binds, depth 10,000 | CE ~1,895× |
The permit, queue, and fiber results make Kyo worth a closer look. Around 5× the throughput for uncontended permits and 3.7× for queue-backed stream batches are encouraging numbers, even for these small workloads.
A little CPU work changed the results too. Bounded workers went from a small Kyo lead to CE ahead by about 2.22×. Kyo’s lead in parallel stream batches shrank to about 1.11×.
Then there’s the last row. CE handled the 10,000-step left-associated bind chain at about 1,895× Kyo’s throughput. Kyo allocated about 1.6 GB per complete chain, versus roughly 943 KB for CE. Making the chain ten times longer increased Kyo’s allocation about a hundredfold. That points to quadratic scaling in this case, and it needs a closer look.
These results make me more interested in trying Kyo. There’s enough here to justify a small application experiment, and a specific weakness to investigate along the way.
This is one machine and a set of small workloads. Cancellation, resource safety, and real I/O still need their own testing.
Source, full results, confidence intervals, and reproduction commands.
If you spot an unfair comparison, please point me to the code. I’d like to get it right.
No comments:
Post a Comment