Vol. I  ·  Data Engineer  ·  Dallas, Texas

Ajaykumar Balakannan

Pipelines, streaming and cloud data platforms.

move the cursor to disturb the field · click to pulse
Data Engineer / Pipelines & Platforms Ajaykumar Balakannan Dallas, Texas

BUILDING WHAT
GETS USED.

A data engineer building the pipelines and platforms everything downstream depends on.

The practitioner

For the past three years I've built the pipelines other people's work sits on top of. Batch and streaming, across healthcare claims, industrial sensor data, job-market data and university research.

At Cigna I run a HIPAA-compliant ELT pipeline over 15TB of claims a day and a Kafka eligibility service answering in under two seconds. Before that, a serverless ingestion layer at Maryland, and a Databricks lakehouse at Siemens that put batch and streaming behind one door for 100+ analysts.

The tools change. What hasn't is the approach: know where the data came from, assume it is wrong until it proves otherwise, and make the thing fail loudly rather than quietly.

A pipeline nobody trusts isn't finished.

Fig. 001 · The author, printed.
Inside this portfolio
Work

Where I've shipped

Sep 2025 - Present
The Cigna Group
Texas, US

Data Engineer

  • Orchestrated a HIPAA-compliant ELT pipeline on AWS Glue and PySpark over 15TB of healthcare claims a day, cutting end-to-end runtime 42% with dynamic partition pruning and reworked join strategies.
  • Built real-time member eligibility validation on Kafka and Spark Structured Streaming, holding sub-two-second latency across 8,500+ concurrent API requests.
  • Deployed a data quality framework with Great Expectations and AWS Deequ, watching 240+ validation rules across Redshift and S3 and lifting SLA adherence to 99.97%.
  • Refactored PySpark jobs from Pandas UDFs to native Spark SQL window functions, dropping compute cost 38% and scaling to 4.2 billion records per batch cycle.
  • Added a Delta Lake layer on Databricks with ACID transactions and time travel, so a failed run rolls back to a point in time instead of being reprocessed from scratch.
AWS GluePySparkKafkaSpark StreamingRedshiftDelta LakeDatabricksGreat Expectations
May 2025 - Aug 2025
Canaria Inc.
New York, NY

Data Science Intern

  • Built Python and SQL/PySpark pipelines over 15M+ job-market records across ClickHouse and PostgreSQL, surfacing trends and data quality problems that fed production decisions.
  • Tuned a Gradient Boosting salary model across 40K+ US zip codes, engineering SOC code, seniority and company features into an auto-updating ClickHouse pipeline.
  • Containerised and orchestrated the pipelines and ML workflows with Docker, Kubernetes and AWS (EC2, S3).
  • Wired LLM APIs (ChatGPT, Claude, Gemini) and text embeddings into compensation data, replacing a manual content workflow with an automated insight-to-content path.
PythonPySparkClickHousePostgreSQLDockerKubernetesAWSLLM APIs
Sep 2024 - May 2025
University of Maryland
College Park, MD

Graduate Data Engineer

  • Created a serverless ingestion pipeline on AWS S3, Glue and Redshift, cutting load latency 40% across 500+ research datasets.
  • Engineered PySpark ETL to cleanse and transform 2TB of raw institutional data, lifting downstream analytics accuracy 25% through stricter validation rules.
  • Untangled workflow dependencies with Apache Airflow, cutting manual intervention and failed retries 60% across 50+ daily DAGs.
  • Tuned Redshift distribution and sort keys, taking 35% off analytical query time for the financial reporting dashboards.
  • Set up completeness checks with Great Expectations, holding 99.5% completeness and catching schema drift before it broke anything downstream.
AWS S3AWS GlueRedshiftPySparkAirflowGreat Expectations
Jan 2023 - Jul 2024
Siemens
India

Data Engineer

  • Delivered a Databricks lakehouse unifying batch and streaming, putting one door in front of the data for 100+ analysts.
  • Tuned Spark SQL queries and partitioning strategy, taking 25% off resource consumption and $3,000 a month off cloud compute.
  • Built a fault-tolerant ETL pipeline in Python and AWS Glue over 10TB of IoT sensor data, holding a 99.9% reliability SLA.
  • Instituted lineage and governance with AWS Lake Formation, keeping 20+ data products ISO 27001 compliant.
DatabricksSpark SQLAWS GlueAirflowLake FormationPython
Toolkit

What I reach for

Pipelines & Orchestration

ETL / ELTApache AirflowApache KafkaSpark Structured StreamingWorkflow OrchestrationBatch & Real-Time ProcessingData Integration

Big Data & Cloud Platforms

Apache SparkSpark SQLPySparkDatabricksDelta LakeAWS (S3, Glue, Redshift, EC2)Azure Data FactoryHadoop

Languages & Data Stores

PythonSQLBashJavaPostgreSQLSnowflakeClickHouseMongoDBSQL ServerPandasNumPy

Quality, Governance & DevOps

Great ExpectationsAWS DeequData ModelingData WarehousingLake FormationDockerKubernetesGitCI/CDJenkins
Projects

Things I built to find out if I could

01

TravelGenie

An agentic trip planner that orchestrates seven MCP servers through Claude, pulling live flights, hotels and events into a single request. I load-tested EC2 against ECS Fargate at 50 to 100 concurrent users; Fargate matched EC2's 460 RPS on half the compute.

Agentic AI · Dec 2025 · Claude / MCP / SerpAPI / AWS
02

Farmland Intrusion Detection

A YOLO vision pipeline that spots wildlife on farmland in real time and fires a non-lethal deterrent matched to the animal, so nobody has to watch a monitor. Hardware prototype built and tested alongside the model.

★ 2nd of 500+ teams · National Technocraft Hackathon · published in BIOGECKO (Vol 12, Iss 03, 2023)
Computer Vision · Mar 2023 · YOLOv3 / CNN / Arduino
03

Bitcoin Price Forecasting

End-to-end pipeline: live BTC collection, ARIMA and XGBoost models at around 85% accuracy, automated ETL, and Qlik Sense dashboards giving real-time visibility into volatility. Cut the manual reporting effort by 40%.

Time Series · 2025 · ARIMA / XGBoost / Qlik Sense
04

Multifamily Real Estate Analytics

A full analytics pipeline over a multifamily portfolio: SQLite to engineered features to five business reports, a seven-sheet Excel workbook and Power BI. Surfaced $2.14M of annual loss-to-lease and a 26-point occupancy gap. The data is synthetic, generated with Faker. This was built to demonstrate the pipeline, not to report on a real client.

Analytics · 2025 · Python / SQL / Power BI

Fig. 003 · Four projects, ordered by how much I learned rather than by date.

Education

Where I learned it

Master's in Data Science

University of Maryland
May 2026 · Maryland
GPA 3.74 / 4.0
Natural Language Processing · Machine Learning · Big Data Systems · Deep Learning

B.Tech, Electrical & Electronics Engineering

Sri Krishna College of Engineering & Technology
Aug 2020 - May 2024 · Coimbatore, India
CGPA 8.4 / 10.0
Microcontrollers · Power System Analysis · Python · ML in Energy Systems
Certifications

Verified credentials

AWS Certified Solutions Architect Associate badge

AWS Certified Solutions Architect - Associate

Amazon Web Services · SAA-C03
Verify on Credly

Microsoft Certified: Power BI Data Analyst Associate

Microsoft · PL-300
Verify on Microsoft Learn

Google Data Analytics Professional Certificate

Google · Coursera · 8 courses
Verify on Coursera
Let's build
something useful.

Open to full-time data engineering and data platform roles. The inbox is always on.