Commit f2d9873
[836] xtable-spark-runtime: thin drop-in Spark bundle for 0.4.x (Spark 3.4 + 3.5) (#843)
* [836] xtable-spark-runtime: thin drop-in Spark bundle (Hudi 0.x)
Ports the xtable-spark-runtime packaging slice onto the Hudi 0.x line
(main-hudi-0x, Hudi 0.14 / Spark 3.4). New thin, relocated bundle that
runs an incremental XTable sync inside a Spark job; engines are provided
by the cluster, never bundled.
Module:
- xtable-spark-runtime_${scala.binary.version}: xtable-core compile;
Spark/Hadoop and the engines (Hudi/Iceberg/Delta) provided. Curated
shade allowlist (xtable modules + guava/protobuf/commons-cli relocated);
avro/parquet/jackson NOT relocated (exchanged with the engines). Thin
~3.7 MB bundle.
- XTableSparkSync: standalone spark-submit entry point (Apache Commons CLI).
- XTableSyncService / TableSyncSpec: build an INCREMENTAL ConversionConfig
and run ConversionController.sync; target metadata path = source data
path (required by Hudi; Iceberg data lives under <basePath>/data).
Hudi 0.14 specifics (vs the Hudi 1.x variant on main):
- No hudi-hadoop-common (that split is 1.x); hudi-common provides FSUtils.
- Engine classpath uses the hudi-spark bundle (its regenerated Avro model
classes link on Avro 1.12, which Iceberg 1.9.2 requires) plus
hudi-java-client for the Hudi target's Java write client, with the raw
hudi-common excluded so the bundle's clean DecimalWrapper wins.
ConversionTargetFactory: make ServiceLoader discovery resilient so a
subset of engines works when others are absent (warn + skip on
LinkageError/ServiceConfigurationError; name-based Delta-Kernel check).
ITXTableSparkRuntimeBundle: spark-submits the shaded jar for one case per
direction across Hudi, Iceberg and Delta (source and target), engines on
a flat classpath, asserting data-equivalence over comparable scalar
columns. Requires a Spark 3.4 SPARK_HOME; skipped otherwise.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [836] xtable-spark-runtime: address review feedback (packaging, licensing, RFC)
Publish the shaded jar as the MAIN artifact with a dependency-reduced POM
(shadedArtifactAttached=false, createDependencyReducedPom=true), matching
iceberg-spark-runtime / hudi-spark-bundle. Resolving the coordinate via a
Maven dependency or --packages now yields the relocated bundle and pulls no
un-relocated transitive deps; --jars is equivalent. The thin main jar +
separate -bundle classifier (which re-introduced the cluster guava clash)
is gone. IT findBundleJar() now picks the shaded main jar.
Pass release/scripts/validate_shaded_license_coverage.sh (the existing
allowlist gate): the shade <includes> must equal the runtime dependency
tree, so jackson / scala-library / log4j-1.2-api - which every Spark
runtime supplies - are declared provided (dropped from the tree, kept off
the shaded jar) and excluded from the IT's flat engine classpath so Spark's
own copies win. avro/parquet stay on the flat classpath (the engine's newer
avro must win over Spark 3.4's).
Bundled-dependency licensing: add META-INF/LICENSE-bundled and
NOTICE-bundled (wired via IncludeResourceTransformer) attributing the only
bundled third-party - guava's closure (Apache-2.0) and commons-cli
(Apache-2.0), plus checker-qual (MIT). Remove the dead protobuf-java
allowlist entry and relocation: protobuf resolves as provided and was never
bundled.
RFC-3: scope v1 to the CLI (XTableSparkSync); mark the config-only
XTableSyncListener on-ramp (and XTableSparkConfig / PlanTargetResolver) as a
deferred follow-up. Fix the activation example to the shipped CLI and
document the required engine Avro version + flat-classpath placement.
XTableSyncService: normalize sourceFormat with Locale.ROOT.
Root pom: exclude the shade-generated dependency-reduced-pom.xml from
spotless (no license header; CI runs clean install).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [836] xtable-spark-runtime: trim duplicated pom comments
Consolidate repeated rationale in the bundle pom (no functional change):
- Tell the hudi-common DecimalWrapper / Avro-1.8-1.9 story once (on the
hudi-spark bundle dep); the hudi-java-client exclusion just points to it.
- State "engine Avro must win on a flat classpath" once (engine-classpath
plugin); the avro dep and excludeGroupIds comments reference it.
- Explain jackson/scala-library are Spark-supplied once.
- Fix a stale "Spark 3.5" reference to Spark 3.4 (this is the Hudi 0.x line).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [836] ci: trigger Maven CI + License Check on main-hudi-0x
Add main-hudi-0x to the push and pull_request branch filters of both
workflows (same change as #841) so CI runs on PRs targeting the Hudi 0.x
release line, including this one. A pull_request workflow's branch filter
takes effect from the PR's own branch, so this makes CI fire on #843.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Address review feedback: guard ServiceLoader.hasNext() and drain spark-submit stdout off-thread
- ConversionTargetFactory: hasNext() resolves provider classes lazily and can
throw ServiceConfigurationError too, so move it inside the existing
ServiceConfigurationError|LinkageError guard alongside next().
- ITXTableSparkRuntimeBundle: drain spark-submit stdout on a background thread
so a hung process is caught by waitFor(10min) instead of blocking forever on
readOutput() reading until EOF.
* [836] ci: point Maven CI + License Check triggers at branch-0.4
main-hudi-0x was deleted and replaced by branch-0.4 (the 0.4.0 release line),
so update the push/pull_request branch filters accordingly.
* [836] ci: validate xtable-spark-runtime bundle against a real Spark distro
ITXTableSparkRuntimeBundle spark-submits the shaded bundle jar and
self-skips unless SPARK_HOME is set, so the main CI never runs it. Add a
path-filtered workflow that installs a matching Spark 3.4 distribution
from the Apache archive, sets SPARK_HOME, and runs the failsafe IT.
* [836] ci: run spark-runtime validation on every branch PR (drop path filter)
The workflow is intended to be a required status check on branch-0.4. A
required check whose workflow is skipped by a path filter never reports,
leaving PRs blocked on a pending check, so run it unconditionally on the
covered branches. The Spark distro is cached, so the added cost is the
reactor build plus the ~20s IT.
* [836] xtable-spark-runtime: address review feedback (CLI validation, unix flags, docs)
- XTableSparkSync: validate --sourceformat/--targets up front and fail fast
(before SparkSession creation) on empty or unsupported values. Split the
allowed sets: sources are Hudi/Iceberg/Delta/Paimon/Parquet, targets are
Hudi/Iceberg/Delta (Paimon/Parquet are read-only, no ConversionTarget).
- XTableSyncService.sourceProviderFor: wire Paimon and Parquet source providers
(previously threw UnsupportedOperationException); engines remain cluster-provided.
- Rename CLI long-opts to unix-style lowercase (--basepath, --sourceformat,
--tablename, --datapath, --partitionspec); update javadoc, IT and RFC example.
- basePathToName -> basePathToTableName: handle "/", trailing slashes and null
by throwing with a "pass --tablename" hint instead of an empty table name.
- Add XTableSparkSyncTest covering table-name derivation and source/target
format validation.
- ConversionTargetFactory: log the skipped provider's error class/message so
operators can distinguish an intentionally-absent engine from a linkage error.
- spark-runtime-validation.yml: add a workflow_dispatch spark_version input and
document why the Spark 3.4 line is pinned (Delta 2.4.0 is Spark-3.4-only).
- RFC-3: add a supported-formats/engine-versions section, "from application
code" (Scala/Java + PySpark) activation examples, and list @vinothchandar as
an approver.
* [836] xtable-spark-runtime: support Spark 3.5 via Delta Kernel
Delta source/target now auto-switch from delta-core to the Spark-free
Delta Kernel implementation on Spark 3.5+ (where the bundled delta-core
does not run), controlled by a single --usedeltakernel toggle that is
auto-enabled by Spark version. Hudi/Iceberg sync was already Spark-free.
Also runs the spark-runtime bundle IT on both Spark 3.4.3 and 3.5.9 via
a CI matrix.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [836] Give Maven CI Build and License Check unique job names
Both workflows used a job named "build", so both reported the same
status-check context and could not be required distinctly. Set unique
job names so branch protection (see #848) can
require each on branch-0.4.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* [836] xtable-spark-runtime: add --datasetconfig for multi-table sync
XTableSparkSync accepts a --datasetconfig YAML (sourceFormat, targetFormats,
datasets[]) to sync multiple tables in one run, mutually exclusive with the
single-table --basepath/--sourceformat/--targets flags. The config is read
through the Spark Hadoop config, so it may live on a local or cloud
(s3/gcs/abfs) path, and reuses the same schema as the RunSync utility.
Parsed with SnakeYAML's SafeConstructor (plain maps/lists only, no arbitrary
type instantiation). SnakeYAML is bundled and relocated to
org.apache.xtable.shaded (like guava/commons-cli) because Spark ships its own
version (1.33 on 3.4, 2.0 on 3.5) that would otherwise clash; its multi-release
classes are filtered out so no un-relocated org.yaml classes remain.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>1 parent 7e273e6 commit f2d9873
14 files changed
Lines changed: 2240 additions & 3 deletions
File tree
- .github/workflows
- rfc/rfc-3
- xtable-core/src/main/java/org/apache/xtable/conversion
- xtable-spark-runtime
- src
- main
- java/org/apache/xtable/spark
- resources/META-INF
- test/java/org/apache/xtable/spark
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
33 | 33 | | |
34 | 34 | | |
35 | 35 | | |
| 36 | + | |
36 | 37 | | |
37 | 38 | | |
38 | 39 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
33 | 33 | | |
34 | 34 | | |
35 | 35 | | |
| 36 | + | |
36 | 37 | | |
37 | 38 | | |
38 | 39 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
54 | 54 | | |
55 | 55 | | |
56 | 56 | | |
| 57 | + | |
57 | 58 | | |
58 | 59 | | |
59 | 60 | | |
| |||
937 | 938 | | |
938 | 939 | | |
939 | 940 | | |
| 941 | + | |
| 942 | + | |
940 | 943 | | |
941 | 944 | | |
942 | 945 | | |
| |||
0 commit comments