zgba 站群
Pre-Release of Polars 2.0

Pre-Release of Polars 2.0

By Ritchie Vink on Wed, 2 Sept 2026

Today we are releasing the first release candidate for Polars 2.0. The definite 2.0 release will land in the following weeks. We don’t aim to make a big feature release of Polars 2.0. In fact we hope it to be a boring experience for you. The reason we bump this major version is that we can get rid of design decisions made in the past that currently block us and then we want to change defaults to more sensible settings that will benefit a greater audience. The biggest default change will be that all LazyFrame queries now will run on the streaming engine. Casual Polars users can therefore expect huge improvements in memory usage and performance. In aggregate we expect the streaming engine to be easily 5x faster.

To help users transition to 2.0, we have posted a full migration guide. This post will cover a few of the highlights.

This is the biggest impact change of 2.0. Calling collect on a LazyFrame will now default to the streaming engine, leading to massive memory and performance improvements on most queries for users. The reason this required a major version bump is that the streaming engine doesn’t guarantee row-order by default for certain operations (join, group_by, unpivot, etc.). If you require observable row-order in those operations, you can opt in to that by setting maintain_order=True.

For users who want to keep using the “in-memory” engine as default, they can do so by setting the engine affinity.

Polars aims to be strict and fail fast. Errors should ideally raise up-front, not 20 minutes into a pipeline. Implicit behavior on data-mismatches should be opt-in, not a default, since those mismatches can hide bugs. This strictness has become even more valuable with the rise of AI-driven development. Agents can validate a query’s structure early by calling collect_schema(), which resolves types and catches schema-level mismatches without materializing any data. This ensures fast feedback for the agents, meaning they can iterate faster. Not all errors can be caught during compilation of the query plan, some depend on data. In these cases Polars defaults to stricter behavior to ensure inconsistencies are caught instead of silently producing different results.

Below are a few examples where Polars has gotten more strict:

If you run an is_in expression on different data-types, Polars used to cast both types to their common supertype, even if that conversion was lossy Below is an example with user-ids that can go wrong by silent data-type mismatches.

Before 2.0, user_id gets coerced to Float64 to match flagged_ids. But 9007199254740993 sits above 2^53 (9007199254740992), the largest integer float64 can represent exactly, so it silently rounds down to 9007199254740992.0, giving a false positive.

In 2.0 this raises: InvalidOperationError: ‘is_in’ cannot check for Int64 values in List(Float64) data., users should explicitly cast to deal with lossy type conversion.

Horizontal concat will now check lengths instead of silently filling with null.

In 2.0 this will raise with:

If padding is what you wanted, you have to explicitly opt-in to that with how=“horizontal_extend”. Making that intention clear to the reader.

Another one worth mentioning is the removal of many casts that were ambiguous or should be applied via their dedicated parsing expression, leading to one obvious way to parse data.

Use instead: .cat.to(dtype) for int → categorical, .cat.physical() for categorical → int.

Use instead: .str.to_date() / .str.to_datetime(). These allow you to apply a parsing format, giving you more control over how the data is parsed.

These were just a few examples, but we landed many more strictness improvements. See them all in the migration guide.

We put a lot of effort into making sure you as user or your agent can continue if you used old parameters that are not supported anymore. We added two new typed exceptions for this; polars.exceptions.AttributeRemovedError and polars.exceptions.ArgumentRemovedError that handle removed attributes and methods and removed parameters respectively.

The error messages should point you to the new API instead. Below we show two examples.

Most of the removed functionality has been deprecated for a long time and hopefully should not have affected your pipelines if you have stayed up to date. Reach out to us if you think we should have kept some functionality you relied on.

Polars 2.0 is about better defaults (most importantly the streaming engine) and a better API. We hope this release is rather boring. We don’t gate new features behind major version bumps as we ship them as soon as their ready.

Don’t be mistaken, Polars 2.x will be much better than 1.x. There is a lot in flight that we haven’t talked publicly enough: proper out-of-core support for the streaming engine, a new IO-plugin design, what we think will be the fastest S3 reader out there, major SQL coverage improvements, a cost-based planner, join reordering, and the removal of mmap, which will make our pipelines fully async end to end.

Try the release candidate by installing pip install polars==2.0rc1. Give it a spin and reach out to us here: https://github.com/pola-rs/polars/issues or contact us on discord: https://discord.gg/4UfP5cfBE7.

To stay up to date and receive our latest BETA news, make sure to sign up. And no worries, we won’t spam your inbox.

View original article