Repository navigation
Arrow PyCapsule stream, pyarrow-free to_polars and Excel-compatible XLSB integers (v5.1.1) - #134
Merged
Merged
Conversation
…-api-v5.1.0 chore: promote public API for v5.1.0
…rrow_stream method in Workbook
Bumps OfficeIMO.Excel from 3.4.3 to 3.4.4 --- updated-dependencies: - dependency-name: OfficeIMO.Excel dependency-version: 3.4.4 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: benchmarks ... Signed-off-by: dependabot[bot] <support@github.com>
…d corresponding tests
…ts/ExcelReader.Benchmarks/develop/benchmarks-f3f134d891 Bump the benchmarks group with 1 update
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #134 +/- ##
==========================================
- Coverage 89.47% 89.45% -0.03%
==========================================
Files 181 181
Lines 13611 13612 +1
Branches 2514 2515 +1
==========================================
- Hits 12179 12177 -2
- Misses 954 956 +2
- Partials 478 479 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Benchmark ResultsBaseline: master Moved by at least 10%
46 within ±10%. Full results per groupExcelReader.Benchmarks.ArrowConversionBenchmark
ExcelReader.Benchmarks.ChunkedParseBenchmark
ExcelReader.Benchmarks.ColdStartBenchmark
ExcelReader.Benchmarks.CsvParallelParseBenchmark
ExcelReader.Benchmarks.CsvParallelVsSepBenchmark
ExcelReader.Benchmarks.CsvParseBenchmark
ExcelReader.Benchmarks.CsvReadBenchmark
ExcelReader.Benchmarks.CsvWriteBenchmark
ExcelReader.Benchmarks.DataReaderBenchmark
ExcelReader.Benchmarks.EncryptedWorkbookBenchmark
ExcelReader.Benchmarks.NativeRowReadBenchmark
ExcelReader.Benchmarks.NativeTypedParseBenchmark
ExcelReader.Benchmarks.ParseBenchmark
ExcelReader.Benchmarks.ReadBenchmark
ExcelReader.Benchmarks.RealDataReadBenchmark
ExcelReader.Benchmarks.RealDataTypedParseBenchmark
ExcelReader.Benchmarks.RecordWriteBenchmark
ExcelReader.Benchmarks.StringHeavyReadBenchmark
ExcelReader.Benchmarks.WriteBenchmark
ExcelReader.Benchmarks.WritePathBenchmark
ExcelReader.Benchmarks.XlsReadBenchmark
ExcelReader.Benchmarks.XlsWriteBenchmark
ExcelReader.Benchmarks.XlsxSharedStringHotPathBenchmark
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Workbook.to_arrow_stream(schema, header_row=1, batch_size=10000)returns an
ArrowStreamimplementing__arrow_c_stream__, built with ctypes alone. Any consumer ofthe Arrow PyCapsule interface takes it directly:
pl.DataFrame(stream),pyarrow.RecordBatchReader.from_stream(stream),pd.DataFrame.from_arrow(stream)(pandas 3).The stream is single-use, and an unconsumed stream or a dropped capsule releases the native read.
to_polars()no longer needs pyarrow. Withparallelism=1(the default) it goes throughto_arrow_stream()andpolars.DataFrame, still streamed and still chunked.parallelismotherthan 1 keeps the whole-sheet path through
to_record_batch(), which still requires pyarrow.BrtCellRk'sfloat form whenever it is exact (every integer up to ±524,288, so every date serial), which is what
Excel does. Before, it always used the int form. That is valid MS-XLSB, but calamine ignores a date
style on it, so
polars.read_excelread our XLSB dates back as integers. Same 4-byte record; largerintegers still use the int form or
BrtCellReal.python/benchmarks/bench_polars_engine.pycompares excelreader againstpolars.read_excel(engine="calamine")on plain XLSX/XLSB, and against msoffcrypto-tool + calamine onencrypted files. It uses a fixed schema on both sides, checks the frames are equal before timing,
runs each case in a fresh subprocess and reports peak RSS. On 1M rows × 7 columns excelreader is
3.0–5.3x faster with 4–6x lower peak RSS (Windows, x64).
Housekeeping
rust/excelreader/build.rsfindslib.exethroughcc::windows_registry, so building the importlibrary no longer needs a VS developer prompt (
ccadded as a build dependency).rust/.cargo/config.tomlpointsEXCELREADER_NATIVE_LIB_DIRat the Python package's_lib.PublicAPI.Shipped.txt.Compatibility
XL_ABI_VERSIONstays 5, and there are no C ABI changes.to_polars()returns the same frame as before.that ignore the date style on int-form RK cells (calamine) benefit, and only for newly written files.
Test plan
dotnet test tests/ExcelReader.Tests: 2450 passed. Adds a theory pinning the RK form for0, ±1, 45000 (float) and 524289, 536870911 (int).
pytest python/testsagainst a freshly built native library: 128 passed, 1 xfailed (pre-existing).Adds tests for
pl.DataFrame(to_arrow_stream()), single use, release when unconsumed, release ona dropped capsule, and negative
batch_size.reads those dates correctly.
bench_polars_engine.py --rows 1000000 --n 5against excelreader-native 5.1.0; all cases verified equal.