☰
Apache Spark读写OpenLake:S3A兼容对象存储集成完整教程
2026/10/3 12:43:38 网站建设 项目流程

Apache Spark读写OpenLake:S3A兼容对象存储集成完整教程

【免费下载链接】openlakeOpenLake is a high performance storage engine for efficient LLM inference and GPU Training项目地址: https://gitcode.com/gh_mirrors/ope/openlake

OpenLake 是一个用 Rust 构建的高性能存储引擎,面向 LLM 推理与 GPU 训练场景,内置 S3 兼容对象存储接口。本教程将手把手教你把Apache Spark接入 OpenLake:通过 Hadoop S3A 连接器读写数据,从构建服务器、创建 Bucket,到写入 Parquet 数据集并读回验证,5 个步骤即可完成集成。

为什么把 Spark 数据湖建立在 OpenLake 上?

很多团队用 MinIO 或 RustFS 这类 S3 兼容存储作为数据湖底座,但高并发下吞吐和延迟会明显掉档。OpenLake 基于 Linuxio_uring构建,单毫秒内可达百万级 IOPS,在吞吐与延迟对比中全面领先:

得益于 S3 兼容 API,Spark 无需任何代码改动——只要配置 S3A 端点、密钥和 Path-Style 访问,现有spark.read.parquet("s3a://...")代码即可直接跑在 OpenLake 上。

环境准备清单

开始之前,请确认机器上具备以下条件(步骤基于 Linux + Spark 3.5.5 Docker 验证):

工具用途说明
Rust编译 OpenLake项目固定工具链 1.91.1,见 rust-toolchain.toml
Docker运行 PySpark 容器已启动
curl创建 S3 桶需支持--aws-sigv4签名
Git拉取源码—

环境搭建细节可参考官方开发者文档 docs/developer/environment_setup.rst。

第一步:构建并启动 OpenLake 服务器

先拉取源码并编译出openlaked服务端二进制:

git clone https://gitcode.com/gh_mirrors/ope/openlake cd openlake cargo build --release --workspace

准备两个本地数据目录,并创建单节点配置node0.toml(S3 端口 9000,RPC 端口 9100):

mkdir -p /tmp/openlake-data0 /tmp/openlake-data1
self_id = 0 data_dirs = ["/tmp/openlake-data0", "/tmp/openlake-data1"] s3_addr = "0.0.0.0:9000" rpc_addr = "127.0.0.1:9100" set_drive_count = 2 default_parity_count = 1 region = "us-east-1" [[credentials]] access_key = "openlakeadmin" secret_key = "openlakesecret" [[nodes]] id = 0 rpc_addr = "127.0.0.1:9100" disk_count = 2

💡 也可以直接使用仓库内置的本地配置 storage-tcp-local.toml 并替换数据目录。

启动服务器:

RUST_LOG=info ./target/release/openlaked --config node0.toml

看到 S3 listener 绑定0.0.0.0:9000、cluster bootstrap complete即代表启动成功,保持终端运行:

第二步:验证端点并创建 S3 桶

先用未签名请求确认端点可达(返回 403AccessDenied属正常):

curl --max-time 10 -s --show-error --write-out "HTTP status: %{http_code}\n" \ --output /tmp/openlake-endpoint-response.xml http://127.0.0.1:9000/

再用 AWS SigV4 签名的 PUT 请求创建test-bucket(成功返回 200):

curl --max-time 20 -s --show-error \ --write-out "HTTP status: %{http_code}\n" \ --aws-sigv4 "aws:amz:us-east-1:s3" \ --user "openlakeadmin:openlakesecret" \ --request PUT http://127.0.0.1:9000/test-bucket

第三步:启动 PySpark 会话并接入 S3A

使用官方 Spark 镜像,启动时通过--packages加载 Hadoop AWS 依赖,并用spark.hadoop.fs.s3a.*系列配置把 Spark 指向 OpenLake。关键配置与node0.toml严格对应:

  • fs.s3a.endpoint→ OpenLake S3 端点(容器内访问宿主机用host.docker.internal:9000)
  • fs.s3a.access.key/fs.s3a.secret.key→ 配置中的access_key/secret_key
  • fs.s3a.path.style.access=true→ 自定义 S3 兼容端点必须启用 Path-Style
  • fs.s3a.connection.ssl.enabled=false→ 本地开发用 HTTP
docker run -it --rm \ --name spark-openlake \ -v /tmp:/tmp \ apache/spark:3.5.5 \ /opt/spark/bin/pyspark \ --conf spark.jars.ivy=/tmp/.ivy2 \ --packages org.apache.hadoop:hadoop-aws:3.3.4 \ --conf spark.hadoop.fs.s3a.endpoint=http://host.docker.internal:9000 \ --conf spark.hadoop.fs.s3a.access.key=openlakeadmin \ --conf spark.hadoop.fs.s3a.secret.key=openlakesecret \ --conf spark.hadoop.fs.s3a.path.style.access=true \ --conf spark.hadoop.fs.s3a.connection.ssl.enabled=false \ --conf spark.hadoop.fs.s3a.aws.credentials.provider=org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider

⏳ 首次运行需要几分钟下载 Hadoop AWS 依赖,耐心等待即可。看到SparkSession available as 'spark'后进入下一步。

第四步:用 Spark 写入 Parquet 数据

在 PySpark 提示符下创建一个小 DataFrame 并写入 OpenLake:

df = spark.createDataFrame( [(1, "Alice"), (2, "Bob"), (3, "Charlie")], ["id", "name"], ) df.coalesce(1).write.mode("overwrite").parquet("s3a://test-bucket/users")

命令无报错并返回提示符,即表示数据已成功持久化到 OpenLake 对象存储。

第五步:读回数据验证集成

通过 S3A 路径把刚才的数据读回:

read_df = spark.read.parquet("s3a://test-bucket/users") read_df.show()

输出的三行数据与写入完全一致,说明 Apache Spark 与 OpenLake 的 S3 兼容对象存储集成验证成功:

常见问题排查清单

现象排查方向
建桶失败先用未签名 curl 确认 9000 端口可达,再重跑签名请求
Spark 启动慢首次运行在下载 Hadoop AWS 依赖,仅发生一次
读写报连接错误核对fs.s3a.endpoint与node0.toml中s3_addr是否一致
认证失败确认 access key / secret key 与[[credentials]]完全一致
会话中途失败确认openlaked服务器终端仍在运行

延伸阅读

  • 完整官方教程:docs/examples/spark_openlake.rst
  • 服务器启动日志示例:docs/openlake-server-logs.png
  • S3 兼容 API 实现:crates/openlake_server/src/s3/
  • 集群运维操作:docs/cluster_operations.rst
  • 项目总览:README.md

集成完成后,你可以尝试更大的数据集、其他 Spark 数据源(CSV、Delta 等),或把 OpenLake 作为多节点 GPU 集群的统一对象存储底座。🚀

【免费下载链接】openlakeOpenLake is a high performance storage engine for efficient LLM inference and GPU Training项目地址: https://gitcode.com/gh_mirrors/ope/openlake

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询