The next project seeks to make a demo of streaming data ingestion and processing.
The following tools will be used:
- Postgres v16
- Debezium for reading events in Postgres -Spark Streaming
- Databricks with AWS ec2 (Compute engine) and S3 (Data Lake) -Docker
For this purpose and to facilitate the configuration of all services, we use docker-compose
You will need to copy this configurations to prepare Postgres with debezium
Copy this file the container
docker cp ./pg_hba.conf <postgres-container-name>:/var/lib/postgresql/data/pg_hba.conf
Maybe you will need to restart the postgres and debezium container
We will need to create the tables in the example database.
We are going to configure three sales tables: products customers sales
We will need to connect debenzium and kafka, and create topics to store the data stored here
We are going to create fictitious data for both customers, products and sales.
For products and sales we will do it only once, but for sales you can run it as often as you want
We are going to ingest data from kafka topics, and that same time window will be processed in memory. The writing will be complete for the moment.
