Repository navigation
[SPARK-60101][SQL] Support native data sources implemented in Rust or C++ - #59313
HyukjinKwon wants to merge 11 commits into
Conversation
…(Rust, C++) data sources This adds the columnar data source API, a simpler way to plug a data source into Spark than implementing the Data Source V2 interfaces directly. It follows the Python Data Source API, but data is exchanged only as columnar batches: readers return ColumnarBatches, and writers receive Arrow-backed ColumnarBatches. Spark adapts a data source to Data Source V2, with batch and micro-batch reads, batch and streaming writes, and predicate, column and limit pushdown. A data source is implemented either on the JVM, with the interfaces of org.apache.spark.sql.datasource and a DataSourceProvider, or in native code such as Rust or C++, as a shared library that implements the native methods of NativeBridge as JNI functions and exchanges data through the Arrow C data interface. Spark loads each library into its own copy of NativeBridge, so libraries exporting the same functions can be loaded side by side, and Spark itself ships no native code. A native data source is distributed as a native data source package, a zip file with the extension .sparkpkg that contains a manifest and the library for each platform. Packages are added with spark.addArtifact, in Scala and Python, classic and Spark Connect, or found under spark.sql.dataSource.native.paths; executors get them through the artifacts of the session. The change also adds the spark.sql.dataSource.native.enabled and spark.sql.dataSource.native.paths configurations, the native data source error conditions, the arrow-c-data dependency, and a documentation page with examples in Scala, Rust and C++. Co-authored-by: Isaac <no-reply@databricks.com>
Downstream data sources are only native, so the JVM interfaces of the columnar data source API (DataSource, DataSourceReader, DataSourceWriter, DataSourceStreamReader, DataSourceStreamWriter and DataSourceProvider) are not public anymore. They become the internal ColumnarDataSource traits that the Data Source V2 adapter and the native data sources share, and the Javadoc of NativeBridge now specifies the whole interface on its own. The documentation page is renamed to "Native Data Sources". Co-authored-by: Isaac <no-reply@databricks.com>
A table created with CREATE TABLE ... USING ... OPTIONS (...) passes its options to TableProvider.getTable, while its scans and writes only get the options of their query. The columnar table now merges the options of the table under the options of each scan and write, so that SELECT and INSERT INTO on such a table use them. Co-authored-by: Isaac <no-reply@databricks.com>
…cally Like the Python data sources installed in the Python path, the native data source packages in the native-datasources directory of SPARK_HOME are now found without any configuration. The packages of a session, added with spark.addArtifact or under spark.sql.dataSource.native.paths, take precedence over the installed ones, like the Python data sources registered at runtime take precedence over the installed ones. Co-authored-by: Isaac <no-reply@databricks.com>
- Exclude the `org.apache.spark.sql.datasource` package, which has no
counterpart in the Spark Connect client, from the client compatibility
check, like the other packages specific to classic Spark.
- Escape the Rust and C++ examples of the native data sources page from
Liquid, which failed to parse `{{0, end / 2}, ...}` in the C++ example.
Co-authored-by: Isaac <no-reply@databricks.com>
Instead of the native-datasources directory of SPARK_HOME, find the library of a native data source where native libraries are usually installed, by its file name: the library spark_datasource_<name>, such as libspark_datasource_<name>.so on Linux, implements the data source <name>. Spark looks for it in the native library path: the directories of java.library.path, which include LD_LIBRARY_PATH on Linux, and then the lib directory of each installation prefix in PATH, such as /usr/local/lib for /usr/local/bin. The executors load the library from the same path as the driver, or else find it in their own native library path. The packages of a session still take precedence. Also: - Keep the error of a library that is loaded but rejected, because it cannot be loaded again in another class loader: using it again failed with "already loaded in another classloader". - Add NativeDataSourceConnectSuite, which tests native data sources with a Spark Connect client and server, and move the helpers that build the test library to NativeDataSourceTestUtils. Co-authored-by: Isaac <no-reply@databricks.com>
The Scala linter checks the format of the Spark Connect modules with scalafmt. Co-authored-by: Isaac <no-reply@databricks.com>
|
Just out of curiosity, was the FFM API considered instead of JNI? |
|
@pan3793 Good question. Yes, but FFM cannot be the only bridge yet because of the Java baseline. Spark compiles with FFM would fit this design well, though, and the design leaves room for it:
|
I suppose we can raise our JDK baseline in the upcoming 5.0.0 (early 2027)? And this can also be a 25-specific feature before that happens, with the decoupled Connect client-server architecture, upgrading server-side JDK version should be much easier.
Yeah, this is the biggest advantage in my mind, non-JVM developers see a clean C header file without needing to learn JNI stuff. |
|
I took a look about FFM (and did prototype). I think it's sort of good and bad. For example, we have to release header file together whereas JNI one does not require it. BTW, we will release 4.4.0 in Dec. and 4.5.0 around Mar. since we will do the quarterly release. I think we can keep JNI version, and do FFM one later around Spark 5.0.0. |
for JNI, we define the interface via a Java class ( I suppose, from the perspective of the native datasource connector developer, the latter should be more friendly (hygienic, and save one step to generate it from Java). If we choose to define a header file as the interface directly, we can:
and once we raise our JDK baseline to 25+, we can get rid of the JNI glue code, the ABI keeps no change from the beginning. |
|
Yeah. Let's actually do that. We can deprecate this in 4.5.0, and drop it in 5.0.0 with replacing it to FFM. I am thinking about bringing more native code into Spark 5 in fact .. |
There was a problem hiding this comment.
Thank you for this proposal, @HyukjinKwon. This is great! I reviewed the lookup, distribution and lifecycle of the native data sources, and left 10 inline comments, summarized below in the same order.
Main issues
- One invalid package breaks every lookup miss (
NativeDataSourceRegistry.scalaL65) - Streaming tasks look up the package under the wrong session UUID (
NativeDataSourceRegistry.scalaL168)
Other issues
- Classic
addArtifact(..., file=True)doesn't register the package (session.pyL2299) - Spark comparison semantics (NaN, collation) for pushed predicates (
NativePredicates.scalaL79) - Artifact dedup by file name can ship a different package (
NativeDataSourceRegistry.scalaL152) - The stream writer handle is never closed explicitly (
NativeDataSource.scalaL294) - A failed
commitDataWriterskips the task abort (NativeDataSource.scalaL484) - The registry lookup runs twice for each resolution (
NativeDataSourceV2.scalaL46) - Installed native libraries now block Python data source registration (
DataSource.scalaL686) - The scan description shows the class name instead of the data source name (
ColumnarTable.scalaL148)
The first two items seem to be the most important: one invalid package breaks unrelated lookups, and streaming queries on a cluster with isolated sessions would not find the package on the executors.
…d lifecycle - Skip and log invalid packages during the lookup, unless the manifest of the invalid package lists the requested data source. - Distribute a package when its reader or writer is created, in the active session, which is the cloned session of a streaming query; executors also check the copies in the files of the other sessions by checksum. - Route .sparkpkg files to the JVM in classic addArtifacts regardless of the flags, also in mixed lists, as Spark Connect does. - Do not push predicates on columns with non-binary collated strings, and document the Spark SQL semantics (NaN, -0.0, nulls) of fully evaluated predicates in NativeBridge. - Report a clear conflict when a different package with the same file name is already an artifact of the session. - Close the previous stream writer of a streaming query when it creates a new one, and document the stream writer lifecycle. - Keep the data writer handle when commitDataWriter fails, so that Spark aborts it with abortDataWriter. - Reuse the lookup of DataSource.lookupDataSource in NativeDataSourceV2. - Skip native data sources when registering a Python data source, which takes precedence over them. - Describe the scan by the name of the data source. Co-authored-by: Isaac <no-reply@databricks.com>
dongjoon-hyun
left a comment
There was a problem hiding this comment.
Thank you for addressing all the comments, @HyukjinKwon. I checked 9dd4f88 against each thread, and all 10 points are handled. I left 4 more minor, non-blocking comments:
- An invalid package that lists the name hides a valid one (
NativeDataSourceRegistry.scalaL88) - The thread-local can keep a stale lookup on pooled threads (
NativeDataSourceRegistry.scalaL70) - The package cache is never evicted (
NativeDataSourcePackage.scalaL102) - No end-to-end test for closing the previous stream writer (
NativeDataSource.scalaL565)
sarutak
left a comment
There was a problem hiding this comment.
I went over the lookup/distribution and the reader/writer lifecycle again on 9dd4f88. The fixes for the earlier round all check out in the code, and the overall design (per-library BridgeClassLoader for side-by-side loading, NativeHandle backed by AtomicLong plus Cleaner, Arrow struct release/close on every path, tryWithSafeFinally around the native calls) is solid.
I left four inline comments, none blocking:
distributecollapses "no active session" and "local mode" into one branch, so a missing active session would silently skip distribution rather than fail clearly. I could not find a planning path that actually has no active session, so this is defensive hardening.- The batch writer's
commitcloses the handle outsidetryWithSafeFinally, unlikeabort; iflibrary.committhrows, the driver-side handle falls back to theCleaner. One-line symmetry fix. spark.sql.dataSource.native.enableddefaults totrue: a note on whether opt-in is preferable for the first release of a feature that loads native code (nit).NativeLibrary.loadusessetAccessibleacross class loaders: a comment on why would help future readers (nit).
The last two are nits. Thanks for the very thorough design and docs on this.
…m writers - Prefer a usable package or an installed library over an invalid package or a package of an unsupported interface version that lists the same data source, and only report those when nothing else provides it. - Clear the lookup kept for the provider at the start of every DataSource.lookupDataSource, so a provider never takes an earlier lookup. - Bound the cache of read packages, and forget the packages of a session when its artifacts are cleaned up. - Record writer closes in the test library and check end to end that each micro-batch closes the stream writer of the previous one, and that each micro-batch creates its stream writer with the ID of the query. Co-authored-by: Isaac <no-reply@databricks.com>
…batching - Resolve the provider of classic DataFrameWriter commands and of DataStreamWriter sinks with the session of the Dataset active, so that the lookups of Python and native data sources use that session. - Distribute a package that is not an artifact of the session as a copy named after its file name and checksum, which the lookups do not count as a package of the session: a configured package can be replaced in place, and packages with the same file name no longer collide. Remove the then unreachable NATIVE_DATA_SOURCE_PACKAGE_FILE_NAME_CONFLICT error. - Log a warning when a package is distributed without an active session. - Skip, and log, a configured directory that cannot be read. - Compare the data with the read schema as it round-trips through Arrow, which does not carry collations, interval fields and user-defined types. - Also end the write batches at spark.sql.execution.arrow.maxBytesPerBatch. - Document that a failed batch commit is aborted, which closes the writer, why NativeBridge.load is called reflectively, and the security of discovering libraries by file name. Co-authored-by: Isaac <no-reply@databricks.com>
|
One small thing before merge: could you update the PR description to mention
Since 9ac5a49, batches are also bounded by |
|
BTW, I plan to leave this open a while. I want others to take a look too :-). |
|
@sarutak Thanks, updated. The Internals section now says writes convert rows to Arrow batches "that end at While there, I went through the rest of the description against 9ac5a49 and updated the other parts that changed during the review:
|
LuciferYang
left a comment
There was a problem hiding this comment.
I went through 9ac5a49 again. The fixes from the earlier rounds hold up in the code. Four more inline comments, each checked on 9ac5a49 with a small test suite: Arrow Java for the first two, and a JVM ColumnarDataSource for the last two.
The first one is the one I'd like fixed before merge: when a library returns a sliced Arrow array, Spark reads the wrong rows without any error. The other three are places where checkSchema lets through a layout that ArrowColumnVector can't read, or where NativeBridge promises something Spark doesn't do (calling abort when a write fails, and calling pushLimit at most once).
…h limits once - Read the Arrow stream of a native library array by array, and reject an array with a non-zero offset at any level, which Arrow Java would read from the start of its buffers, with a NATIVE_DATA_SOURCE_ERROR that names the column. - Name the fields of the entries of maps `key` and `value` before creating the vectors, as ArrowColumnVector looks them up by name, so that maps built by arrow-rs can be read; reject decimal256, and report any other column that ArrowColumnVector cannot read with its name. - Create the writer of a batch or streaming write with its writer factory, after the stages that the write depends on ran, so that Spark commits or aborts each writer it creates, also when such a stage fails with adaptive execution. - Ask the reader for a limit only once, as NativeBridge documents, although Spark pushes the limit of LIMIT ... OFFSET twice. - Document these in NativeBridge and the native data sources page. Co-authored-by: Isaac <no-reply@databricks.com>
What changes were proposed in this pull request?
This PR adds native data sources: data sources implemented in native code, such as Rust or C++, without any JVM code. A native data source is a shared library that implements a binary interface. Spark finds it automatically when it is installed like other native libraries, or in a package file added to a session, and loads it when a query uses it. Spark adapts it to Data Source V2, with batch and micro-batch streaming reads, batch and streaming writes, and predicate, column and limit pushdown. Data is exchanged only as columnar batches, through the Arrow C data interface.
Public API: a single class,
org.apache.spark.sql.datasource.NativeBridge(@Evolving, since 4.4.0). A library implements itsnativemethods as JNI functions, such asJava_org_apache_spark_sql_datasource_NativeBridge_createDataSource, and its Javadoc is the specification of the interface: the lifecycle, the handles, the ownership of the Arrow structs, errors (Java exceptions, reported asNATIVE_DATA_SOURCE_ERROR), and the JSON format of the pushed predicates, with the Spark SQL semantics that a predicate evaluated by the library must follow, for example for NaN. Spark loads each library into its own copy of the class, with a separate class loader, so any number of libraries exporting the same functions can be loaded side by side. Spark itself ships no native code for native data sources, besides the JNI library that the newarrow-c-datadependency bundles. The interface uses JNI rather than the FFM API, which is final only from Java 22 while Spark supports Java 17 and 21; the interface is versioned, so that a later version can be a plain C ABI called through FFM.abiVersion,createDataSource,closeDataSourceschemacreateReader,partitions,read,closeReaderpushPredicates,pushLimit,pruneColumns,serializeReadercreateStreamReader,initialOffset,latestOffset,streamPartitions,read,closeStreamReaderserializeStreamReader,commitOffsetcreateWriter,commit,abort,closeWriter,createDataWriter,write,commitDataWriter,abortDataWriterserializeWritercreateStreamWriter,commit,abort,closeWriter,createDataWriter,write,commitDataWriter,abortDataWriterserializeWriterA library only exports the functions of the operations it supports. Spark reports a missing required function as
DATA_SOURCE_*_NOT_SUPPORTED.Finding native data sources: when no Java or Python data source has the requested name, Spark looks for a native data source with that name, in two places:
Installed libraries, automatically: like the Python data sources installed in the Python path, a library installed like other native libraries is found without any configuration, by its file name. The library of the data source
rust_rangeisspark_datasource_rust_range:libspark_datasource_rust_range.soon Linux,libspark_datasource_rust_range.dylibon macOS andspark_datasource_rust_range.dllon Windows. Spark checks whether the file exists in the native library path, which consists of:java.library.path, where the JVM itself finds native libraries. They includeLD_LIBRARY_PATHon Linux,DYLD_LIBRARY_PATHon macOS andPATHon Windows, and sospark.driver.extraLibraryPathandspark.executor.extraLibraryPath;libdirectory of each installation prefix inPATH:<prefix>/libfor each<prefix>/bin.So libraries are found where native libraries are usually installed:
/usr/local/lib(make install, CMake, cargo-c, Homebrew on Intel macs),/opt/homebrew/lib,$CONDA_PREFIX/liband$VIRTUAL_ENV/libin an activated environment, and/usr/lib. A library that implements several data sources is installed under each of their names, for example with symbolic links. The executors load the library from the same path as the driver, or else find it in their own native library path.Packages of a session: a native data source package is a zip file with the extension
.sparkpkg, with a manifest,spark-native-datasource.json, that lists the ABI version, the names of the data sources, and the library for each platform such aslinux-x86_64orosx-aarch64. A session adds packages withspark.addArtifact(Scala and Python, classic and Spark Connect:.sparkpkgfiles are added as file artifacts), or withspark.sql.dataSource.native.paths. They take precedence over the installed libraries, and the executors get them through the artifacts of the session, except in local mode, where they read them in place: a configured package is copied there under a name with its checksum, so it can be replaced in place. Spark ignores, and logs, a package that it cannot read or whose ABI version it does not support, unless its manifest lists the requested data source and no other package or installed library provides it.Short names: no JVM class is needed for each data source. After the Java data sources (
DataSourceRegister) and the Python ones,DataSource.lookupDataSourcelooks the name up in the packages of the session and in the native library path, case-insensitively, and resolves it to a single provider class,NativeDataSourceV2. The caller then gives it the requested name, the wayPythonDataSourceV2already works: a smallNamedTableProvidertrait generalizes howPythonDataSourceV2gets its short name, so the native provider shares the same call sites. The provider reuses the result of the lookup that found its class, instead of looking the name up again. A Python data source can still be registered with the name of a native data source, and takes precedence over it. Lookups only read manifests and check whether files exist; Spark loads a library when a query uses it. Native data sources work withspark.read/write,readStream/writeStream, and SQL tables (CREATE TABLE ... USING rust_range OPTIONS (...),SELECTandINSERT INTO).Internals:
execution.datasources.v2.columnar: an internal columnar data source interface, adapted to Data Source V2. Writes convert rows to Arrow batches that end atspark.sql.execution.arrow.maxRecordsPerBatchrows or once they holdspark.sql.execution.arrow.maxBytesPerBatchbytes, whichever is reached first, as for Python UDFs. The writer of a write is created once the stages it depends on ran, right before its tasks start, so that each writer is committed or aborted. A micro-batch streaming write gets a new stream writer for each micro-batch, and the previous one of the query is closed when the next one is created.execution.datasources.v2.ffi: the discovery of packages and installed libraries, package parsing, the distribution of packages to the executors, library loading, the implementation of the internal interface on top ofNativeBridgewitharrow-c-data, which reads the Arrow streams array by array to reject arrays with a non-zero offset (Arrow Java ignores offsets), and the JSON encoding of pushed predicates.Other changes:
spark.sql.dataSource.native.enabled(static, defaulttrue) andspark.sql.dataSource.native.paths.NATIVE_DATA_SOURCE_ERROR,INVALID_NATIVE_DATA_SOURCE_LIBRARY(CANNOT_LOAD,UNSUPPORTED_ABI_VERSION),INVALID_NATIVE_DATA_SOURCE_PACKAGE(INVALID_MANIFEST,INVALID_LIBRARY,UNSUPPORTED_ABI_VERSION,UNSUPPORTED_PLATFORM),NATIVE_DATA_SOURCE_LIBRARY_NOT_FOUND,NATIVE_DATA_SOURCE_PACKAGE_CONFLICTandNATIVE_DATA_SOURCE_PACKAGE_NOT_FOUND.org.apache.arrow:arrow-c-data, with the same version asarrow-vector. It bundles its own JNI library for Linux, macOS and Windows.DataSource.lookupDataSourcelooks up the native data sources after the Python ones, andDataSourceRegistrationskips them when it registers a Python data source.ArtifactManagerlists the native data source packages of a session, and forgets the cached packages read from the artifacts of a session when it cleans them up.DataFrameWriterbuilds its commands, andDataStreamWriterresolves its sink, with the session of the Dataset active, as the analysis of a query does: the lookups of the Python and native data sources use the active session.addArtifactsadds.sparkpkgfiles to the session whatever the flags: in classic mode through the JVM session, and with Spark Connect as file artifacts.CheckConnectJvmClientCompatibilityexcludes the neworg.apache.spark.sql.datasourcepackage, which the Spark Connect client does not have.Examples
Both examples implement a read-only data source that returns the numbers
[0, end)as a columnid, in two partitions, with the optionend(10 by default). They only export the 8 functions needed for batch reads.Rust (with the
jnicrate and arrow-rs)Cargo.toml:src/lib.rs:Build and install, here on macOS:
Or package, to add it to a session with
spark.addArtifact:C++ (with Arrow C++)
cpp_range.cc:Build (Arrow C++ 24 requires C++20) and install, or package as above:
Usage, in classic PySpark or with Spark Connect. Once installed, the data sources are found automatically, without
spark.addArtifactor any configuration:Or added to a session as a package:
Why are the changes needed?
Many data systems have their best client libraries in native code, such as arrow-rs, DataFusion, delta-rs, iceberg-rust and Lance in Rust, or C++ libraries. Using them from Spark today requires writing a JVM Data Source V2 connector with its own JNI glue. With this change, a native data source implements a stable binary interface once, is installed like any other native library or distributed as a single file, and works with every Spark API, including Spark Connect clients in any language, without any JVM code.
Does this PR introduce any user-facing change?
Yes. It adds native data sources, with the
NativeBridgeinterface, the configurations and the error conditions above. Spark finds the native data source libraries installed in the native library path, andspark.addArtifactnow accepts native data source packages (.sparkpkgfiles). Also, in classic mode, writing with a session that is not the active one now finds the data sources of that session, such as its Python data sources, becauseDataFrameWriterandDataStreamWriterresolve the data source with that session active.How was this patch tested?
NativeDataSourceSuite, with a dependency-free C++ test library (sql/core/src/test/resources/native-datasource/test_native_datasource.cc) compiled by the tests with the C++ compiler of the machine; the tests that load it are skipped if there is no compiler. It covers inferred and user-specified schemas, partitions and batches, the supported types, predicate, column and limit pushdown, batch writes (append and overwrite), aborted writes, streaming reads and writes over several micro-batches with offset commits, SQL tables withOPTIONS(SELECTandINSERT INTO), errors reported by native code, missing functions, several libraries exporting the same functions side by side,spark.addArtifact, packages in directories, package conflicts, the precedence of Java data sources, and invalid packages (manifest, ABI version, platform and library). It also has unit tests for the manifest, the platform names, the native library path and the JSON encoding of predicates.NativeDataSourceSuitealso covers installed libraries: found automatically in the native library path with case-insensitive names, a library installed under several names, writes, the precedence of the packages of a session, the lookup by name on executors (NATIVE_DATA_SOURCE_LIBRARY_NOT_FOUND), and invalid installed libraries (CANNOT_LOAD, andUNSUPPORTED_ABI_VERSION, with the same error when the library is used again).NativeDataSourceConnectSuite, with a Spark Connect client and server:spark.addArtifactof a package, reads, writes, SQL tables, the isolation of the packages of different sessions, and installed libraries.ColumnarDataSourceSuite, which tests the internal Data Source V2 adapter with a data source implemented on the JVM..sparkpkgartifacts in classic mode, whatever the flags and in lists with other files, and with Spark Connect.NativeDataSourceSuiteunless noted:commitDataWriteraborts the data writer; a failedcommitaborts the write, then closes the writer; the test library records when a writer is closed, which checks that a batch write closes its writer, and that each micro-batch of a streaming write closes the writer of the previous one;ColumnarDataSourceSuitechecks that each micro-batch creates its stream writer with the ID of the query, and that batches end atspark.sql.execution.arrow.maxBytesPerBatch.decimal256are rejected with the name of the column, maps whose entry fields are namedkeysandvalues(as arrow-rs names them) are read, andlimit(...).offset(...)asks the library for the limit once.ColumnarDataSourceSuite).libdirectory of an installation prefix inPATHwithout any configuration (and not found without it), found throughjava.library.path, added withspark.addArtifact, and as SQL tables.-Wall -Wextra), andNativeDataSourceSuiteruns on Linux with g++ in CI.ResolvedDataSourceSuite,ArtifactManagerSuite,PythonDataSourceSuite,PythonStreamingDataSourceSuite,DataStreamReaderWriterSuite,DataFrameReaderWriterSuite,DataFrameWriterV2Suite,DataSourceSuite,DDLSourceLoadSuite,DataSourceV2Suite,StreamingDataSourceV2Suite,SQLConfSuite,SparkThrowableSuiteandSparkConfigBindingPolicySuite.Was this patch authored or co-authored using generative AI tooling?
Generated-by: Claude Code (Claude Opus 5.5)
This pull request and its description were written by Isaac.