Internship: Exploring Real-Time Data Processing at Scale

Vor 5 Tagen

Pully, Schweiz ELCA Vollzeit
Description

Start date: February 2027 · Duration: 6 months
Please include your latest academic transcripts with your application. Applications submitted without transcripts will not be considered.

In this 6-month internship starting in February 2027, you will explore real-time data processing at scale with Apache Kafka, Apache Flink and Apache Spark, and help us define guidelines for choosing the right framework.

Your tasks

  • Understand the fundamental concepts of Kafka, Flink and Spark, including their architecture and use cases
  • Implement a pipeline processing streaming data from a single source using Kafka and Flink/Spark, going as far as possible with insights and optimizations
  • Build a second, more complex pipeline: Database → Debezium → Kafka → Flink/Spark → transactional and analytical queries
  • Handle multiple tables and implement watermarking to ensure synchronized data processing
  • Compare Flink and Spark on performance, ease of use and suitability for specific use cases
  • Document your findings and propose guidelines for choosing between the two frameworks

What we offer

  • A dynamic, collaborative workplace with a highly motivated, multicultural team working across international sites
  • The chance to make a difference in people's lives by building innovative solutions
  • Internal coding events such as Hackathons and Brownbag sessions, plus our technical blog
  • Monthly After-Works organized at each location

About your profile

  • Bachelor's or Master's student in Computer Science, Data Engineering or a related field
  • Interest in distributed systems and real-time data processing
  • Programming skills (e.g. Java, Scala or Python) and knowledge of SQL
  • Curiosity, autonomy and good analytical skills