Tomasz Michalski Solutions · Data Engineer

I build systems
that don't wait
for batch.

I design and implement real-time event processing pipelines - from the producer, through Pub/Sub and BigQuery, to ready-to-use data. I operate as an independent contractor on B2B terms.

01 / focus

Data in motion, not at rest

Most systems still treat data as something that arrives once a day. I design architectures that react the moment an event occurs - event streams, window aggregations, and stateful processing.

I operate as an independent contractor and work mainly in the B2B model - whether as an extra pair of hands for a specific pipeline, a streaming architecture consultant, or the person who sets up Kafka and Flink from scratch.

Cooperation model
B2B / Freelance, remote or hybrid
Main focus
event streaming, real-time processing, data backend, GCP
Languages
Python, Go, SQL
Availability
projects and long-term contracts - ask for dates

02 / stack

Tools I use daily

layer tool purpose
event transport Apache Kafka & GCP Pub/Sub queuing, global message routing, event stream ingestion
processing Apache Flink & Cloud Dataflow stateful stream processing, managed Apache Beam jobs
orchestration Cloud Composer (Airflow) workflow scheduling, DAGs, ETL/ELT pipeline automation
data warehouse BigQuery & SQL serverless analytics, massive scale modeling, data sinks
data lake Google Cloud Storage (GCS) object storage, staging areas, raw data landing
services Python, Go lightweight, efficient producers and consumers
environment Docker, Kubernetes reproducible environments, managed container orchestration

03 / projects

Selected projects

Warsaw Transit Data Pipeline

Airflow · dbt · Python · Docker

An end-to-end ETL data pipeline orchestrating the extraction and processing of real-time public transit telemetry.

  • Engineered automated data ingestion using Python and Apache Airflow to extract real-time API data and load it directly into an AWS S3 Data Lake in Parquet format.
  • Implemented serverless analytics and data modeling by leveraging AWS Glue, Athena, and dbt, while provisioning the entire cloud infrastructure via Terraform.
View project →

GitTrends Data Pipeline

Airflow · dbt · Terraform · Docker

A Medallion Architecture-based data pipeline designed to orchestrate the end-to-end extraction and processing of large-scale GitHub event logs.

  • Architected scalable data ingestion and processing using Python and Apache Airflow to load raw data into an AWS S3 Data Lake, performing complex data flattening and transformation with PySpark on Databricks.
  • Implemented serverless analytics and dimensional data modeling by leveraging AWS Glue, Athena, and dbt, while fully provisioning the cloud environment using Terraform.
View project →

04 / services

How we can work together

Streaming pipeline implementation

Design and implementation of event processing from scratch.

Architecture consulting

Review of existing systems, recommendations on scaling, partitioning, and checkpointing.

Batch to streaming migration

Transitioning from periodic processing to event-driven processing, step by step.

Team support / joining a project

Temporary or permanent support for your data team as an additional engineer in the B2B model.

  1. 1

    Chat

    A short conversation about the context, data, and goals - no strings attached.

  2. 2

    Analysis

    Evaluation of the current architecture and a proposed approach, including estimates.

  3. 3

    Implementation

    Iterative work with regular checkpoints.

  4. 4

    Deployment and handover

    Production launch and documentation enabling further work by your team.

05 / contact

Let's talk about your pipeline

The fastest way is via email - briefly describe your project context and the technologies you currently use.

06 / else?

Searching someone else?

Wrong Tomasz? Here are people whose work I can vouch for.