The idea is to connect test results and artifacts with commit history and testing environment. This way test reports can show whether a failure in a pull request is a real regression or a known flake from a target branch.
The core of the service is a specialized analytics engine that makes it possible to store and query large amounts of test results [1]. Thanks to this engine, we've processed 2.5 billion test results to date.
Flakiness.io is used already by a few large open sources (nuxt/nuxt [2], wordpress/gutenberg [3]). There's a 1GB free plan which is enough for roughly 10M+ test results, and I'd be happy to increase this limit to 100GB for large open source projects upon request.
I've been working on shard balancing for Playwright Test. The idea is to use the new API released in Playwright v1.62 together with historical test durations to balance tests across shards.
I ended up designing a specialized variant of LPT that works well for Playwright test suites. It's now used in WordPress's Gutenberg project, where it cut test runtime from 35 minutes to 23.
The algorithm is called SALT. Here's the detailed write-up if you're interested in job scheduling:
The idea is to connect test results and artifacts with commit history and testing environment. This way test reports can show whether a failure in a pull request is a real regression or a known flake from a target branch.
The core of the service is a specialized analytics engine that makes it possible to store and query large amounts of test results [1]. Thanks to this engine, we've processed 2.5 billion test results to date.
Flakiness.io is used already by a few large open sources (nuxt/nuxt [2], wordpress/gutenberg [3]). There's a 1GB free plan which is enough for roughly 10M+ test results, and I'd be happy to increase this limit to 100GB for large open source projects upon request.
Let me know what you think about the project!
[1]: https://blog.flakiness.io/posts/2026/engine/
[2]: https://flakiness.io/nuxt/nuxt
[3]: https://flakiness.io/wordpress/gutenberg