Video by CNCF [Cloud Native Computing Foundation] via YouTube

Don’t miss out! Join us at our next KubeCon + CloudNativeCon events in Shanghai, China (8-9 September, 2026) and Salt Lake City, United States (Nov 9–12, 2026). Connect with our current graduated, incubating, and sandbox projects as the community gathers to further the education and advancement of cloud native computing. Learn more at https://kubecon.io
Interactive Spark at Your Fingertips: Integrating SparkConnect into Kubeflow Notebooks – Vikas Saxena, RAICS.AI
Background Apache Spark is indispensable for large-scale data processing and feature engineering in ML pipelines. However, running Spark interactively within a Kubeflow Notebooks environment has historically relied on tools like Jupyter Enterprise Gateway and Apache Toree — both of which are no longer actively maintained, making them a security risk, a source of operational debt, and unlikely to support future Spark versions. Why SparkConnect (and Almond for Scala users)? SparkConnect is a Spark-native, actively maintained alternative which, combined with the Spark Operator, makes Spark a true native service inside Kubeflow — no third-party gateway, no external process management. Because the connection is interactive, data scientists can run exploratory analysis and quick proof-of-concepts against large datasets without spinning up dedicated jobs — iterating cell by cell, at scale, just like working with a local DataFrame. What we’ll cover A live walkthrough of: Deploying SparkConnect as a Kubernetes service Connecting to the deployed service using PySpark Connecting to the deployed service using the Scala API