Apache Spark读写OpenLake:S3A兼容对象存储集成完整教程
【免费下载链接】openlakeOpenLake is a high performance storage engine for efficient LLM inference and GPU Training项目地址: https://gitcode.com/gh_mirrors/ope/openlake
OpenLake 是一个用 Rust 构建的高性能存储引擎,面向 LLM 推理与 GPU 训练场景,内置 S3 兼容对象存储接口。本教程将手把手教你把Apache Spark接入 OpenLake:通过 Hadoop S3A 连接器读写数据,从构建服务器、创建 Bucket,到写入 Parquet 数据集并读回验证,5 个步骤即可完成集成。
为什么把 Spark 数据湖建立在 OpenLake 上?
很多团队用 MinIO 或 RustFS 这类 S3 兼容存储作为数据湖底座,但高并发下吞吐和延迟会明显掉档。OpenLake 基于 Linuxio_uring构建,单毫秒内可达百万级 IOPS,在吞吐与延迟对比中全面领先:
得益于 S3 兼容 API,Spark 无需任何代码改动——只要配置 S3A 端点、密钥和 Path-Style 访问,现有spark.read.parquet("s3a://...")代码即可直接跑在 OpenLake 上。
环境准备清单
开始之前,请确认机器上具备以下条件(步骤基于 Linux + Spark 3.5.5 Docker 验证):
| 工具 | 用途 | 说明 |
|---|---|---|
| Rust | 编译 OpenLake | 项目固定工具链 1.91.1,见 rust-toolchain.toml |
| Docker | 运行 PySpark 容器 | 已启动 |
| curl | 创建 S3 桶 | 需支持--aws-sigv4签名 |
| Git | 拉取源码 | — |
环境搭建细节可参考官方开发者文档 docs/developer/environment_setup.rst。
第一步:构建并启动 OpenLake 服务器
先拉取源码并编译出openlaked服务端二进制:
git clone https://gitcode.com/gh_mirrors/ope/openlake cd openlake cargo build --release --workspace准备两个本地数据目录,并创建单节点配置node0.toml(S3 端口 9000,RPC 端口 9100):
mkdir -p /tmp/openlake-data0 /tmp/openlake-data1self_id = 0 data_dirs = ["/tmp/openlake-data0", "/tmp/openlake-data1"] s3_addr = "0.0.0.0:9000" rpc_addr = "127.0.0.1:9100" set_drive_count = 2 default_parity_count = 1 region = "us-east-1" [[credentials]] access_key = "openlakeadmin" secret_key = "openlakesecret" [[nodes]] id = 0 rpc_addr = "127.0.0.1:9100" disk_count = 2💡 也可以直接使用仓库内置的本地配置 storage-tcp-local.toml 并替换数据目录。
启动服务器:
RUST_LOG=info ./target/release/openlaked --config node0.toml看到 S3 listener 绑定0.0.0.0:9000、cluster bootstrap complete即代表启动成功,保持终端运行:
第二步:验证端点并创建 S3 桶
先用未签名请求确认端点可达(返回 403AccessDenied属正常):
curl --max-time 10 -s --show-error --write-out "HTTP status: %{http_code}\n" \ --output /tmp/openlake-endpoint-response.xml http://127.0.0.1:9000/再用 AWS SigV4 签名的 PUT 请求创建test-bucket(成功返回 200):
curl --max-time 20 -s --show-error \ --write-out "HTTP status: %{http_code}\n" \ --aws-sigv4 "aws:amz:us-east-1:s3" \ --user "openlakeadmin:openlakesecret" \ --request PUT http://127.0.0.1:9000/test-bucket第三步:启动 PySpark 会话并接入 S3A
使用官方 Spark 镜像,启动时通过--packages加载 Hadoop AWS 依赖,并用spark.hadoop.fs.s3a.*系列配置把 Spark 指向 OpenLake。关键配置与node0.toml严格对应:
fs.s3a.endpoint→ OpenLake S3 端点(容器内访问宿主机用host.docker.internal:9000)fs.s3a.access.key/fs.s3a.secret.key→ 配置中的access_key/secret_keyfs.s3a.path.style.access=true→ 自定义 S3 兼容端点必须启用 Path-Stylefs.s3a.connection.ssl.enabled=false→ 本地开发用 HTTP
docker run -it --rm \ --name spark-openlake \ -v /tmp:/tmp \ apache/spark:3.5.5 \ /opt/spark/bin/pyspark \ --conf spark.jars.ivy=/tmp/.ivy2 \ --packages org.apache.hadoop:hadoop-aws:3.3.4 \ --conf spark.hadoop.fs.s3a.endpoint=http://host.docker.internal:9000 \ --conf spark.hadoop.fs.s3a.access.key=openlakeadmin \ --conf spark.hadoop.fs.s3a.secret.key=openlakesecret \ --conf spark.hadoop.fs.s3a.path.style.access=true \ --conf spark.hadoop.fs.s3a.connection.ssl.enabled=false \ --conf spark.hadoop.fs.s3a.aws.credentials.provider=org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider⏳ 首次运行需要几分钟下载 Hadoop AWS 依赖,耐心等待即可。看到
SparkSession available as 'spark'后进入下一步。
第四步:用 Spark 写入 Parquet 数据
在 PySpark 提示符下创建一个小 DataFrame 并写入 OpenLake:
df = spark.createDataFrame( [(1, "Alice"), (2, "Bob"), (3, "Charlie")], ["id", "name"], ) df.coalesce(1).write.mode("overwrite").parquet("s3a://test-bucket/users")命令无报错并返回提示符,即表示数据已成功持久化到 OpenLake 对象存储。
第五步:读回数据验证集成
通过 S3A 路径把刚才的数据读回:
read_df = spark.read.parquet("s3a://test-bucket/users") read_df.show()输出的三行数据与写入完全一致,说明 Apache Spark 与 OpenLake 的 S3 兼容对象存储集成验证成功:
常见问题排查清单
| 现象 | 排查方向 |
|---|---|
| 建桶失败 | 先用未签名 curl 确认 9000 端口可达,再重跑签名请求 |
| Spark 启动慢 | 首次运行在下载 Hadoop AWS 依赖,仅发生一次 |
| 读写报连接错误 | 核对fs.s3a.endpoint与node0.toml中s3_addr是否一致 |
| 认证失败 | 确认 access key / secret key 与[[credentials]]完全一致 |
| 会话中途失败 | 确认openlaked服务器终端仍在运行 |
延伸阅读
- 完整官方教程:docs/examples/spark_openlake.rst
- 服务器启动日志示例:docs/openlake-server-logs.png
- S3 兼容 API 实现:crates/openlake_server/src/s3/
- 集群运维操作:docs/cluster_operations.rst
- 项目总览:README.md
集成完成后,你可以尝试更大的数据集、其他 Spark 数据源(CSV、Delta 等),或把 OpenLake 作为多节点 GPU 集群的统一对象存储底座。🚀
【免费下载链接】openlakeOpenLake is a high performance storage engine for efficient LLM inference and GPU Training项目地址: https://gitcode.com/gh_mirrors/ope/openlake
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考