Distributed stream processing
2026
SDE Spark
Synopsis Data Engine: a distributed system that maintains and queries probabilistic data structures over continuous Kafka streams, built on Apache Spark Structured Streaming.
- ▸ Hybrid pipeline that cuts stateful operators from 3 to 1, removing two thirds of HDFS checkpoint writes per micro-batch.
- ▸ Stateless hash routing across N slots; partial estimates are merged within the same micro-batch with no cross-batch buffering.
- ▸ Count-Min, Bloom filter, AMS and HyperLogLog synopses, with unit, end-to-end and load-test harnesses.
Pipeline
-
L1 KafkaIngestionLayer
JSON → typed events
-
L2 StatelessRouter
hash-route to N slots
-
L3 SynopsisProcessor stateful
add / estimate / delete
-
L4 PathSplitter
single vs. partitioned
-
L5 ReduceAggregator
merge N partials
-
L6 KafkaOutputLayer
estimations → Kafka