A collection of big data processing and analytics projects built with Apache Spark, Kafka, Flume, Oozie, and Apache Phoenix.
The repository includes batch-processing, stream-processing, ingestion, workflow orchestration, and data-query examples using the MovieLens ml-100k dataset and raw web logs.
- Apache Spark
- Spark Streaming
- Apache Kafka
- Apache Flume
- Apache Oozie
- Apache Phoenix
- Hadoop ecosystem
big-data-analytics/
├── Apache-Phoenix/
├── Kafka/
├── Spark/
├── Spark-Streaming/
├── Flume/
├── Oozie/
├── ml-100k/
└── README.md