Skip to content

bulkcopy_arrow rejects polars DataFrames: Utf8View string columns unsupported #708

Description

@oscarrobertson

Describe the bug

The bulkcopy_arrow docstring says any object exposing arrow_c_stream works as a source, and names polars DataFrames as an example. In practice any polars ≥ 1.0 DataFrame with a string column fails.

Polars stores strings natively in Arrow's string-view layout, so its zero-copy arrow_c_stream export produces Utf8View columns, which the writer doesn't support. Frames without string columns work fine.

Workaround: cursor.bulkcopy_arrow("dbo.T", df.to_arrow()) — polars re-encodes strings to large_string for compatibility. That costs an extra copy of all string data, which is the overhead this API otherwise avoids.

Worth supporting Utf8View (and presumably BinaryView) in the writer — pandas is heading the same way with Arrow-backed strings, so view types will increasingly be what arrow_c_stream hands you. Failing that, the docstring shouldn't cite polars as a working example.

Exception message:

ValueError: Cannot map Arrow column 'name' (Utf8View) to SQL column 'name' (VarChar):
Usage Error: type combination is not supported by the Arrow row-major writer

To reproduce

import polars as pl

df = pl.DataFrame({"id": pl.Series([1, 2], dtype=pl.Int32), "name": ["a", "b"]})
cursor.bulkcopy_arrow("dbo.T", df)  # T: (id INT, name VARCHAR(50))

Expected behavior

That the above code just works.

Further technical details

mssql-python 1.13.0, polars 1.39.3, pyarrow 23.0.1, Linux x86_64.

Metadata

Metadata

Assignees

Labels

triage neededFor new issues, not triaged yet.

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions