Data Scientist · Dallas, TX

Ajaykumar Balakannan

Production ML, NLP & LLM-orchestrated systems. Move your cursor.

drag to orbit · click to pulse
MS Data Science · University of Maryland

Ajaykumar
Balakannan

A
Ajaykumar Balakannan
Status
Open to data
& ML roles
Based in
Dallas, TXopen to relocate / remote
Currently

I build the data systems that help people make better calls: no-show models for a campus counseling center, salary forecasts across 40K+ zip codes, and lately agentic AI that orchestrates Claude and MCP. Production ML and NLP, with a stubborn habit of shipping things stakeholders actually use.

Grad · May 2026
3.74
GPA at UMD
k-means, settling
In production
15M+
records processed
Work

Where I've shipped

Sep 2024 – Present
University of Maryland
College Park, MD

Data Scientist

  • Built a no-show prediction model (Random Forest + LightGBM, SHAP for the "why") on counseling appointment records. It triggers automated outreach and lifted resource utilization by 25%.
  • Classified 4,000+ pieces of student feedback (e-forms and audio transcripts) with SetFit and BERTopic, surfacing sentiment and service themes the counseling staff now act on.
  • Wrote the Python / DuckDB / SQL pipelines that fold 12+ years of operational records into version-controlled data models, the backbone for the models and KPI reporting.
  • Automated the refreshes behind the Tableau and Power BI dashboards (utilization, cancellations, peak demand), cutting manual reporting turnaround by 60%.
PythonDuckDBSQLLightGBMSHAPSetFitBERTopicTableauPower BI
May 2025 – Aug 2025
Canaria Inc.
New York, NY

Data Science Intern

  • Trained a LightGBM salary model across 40K+ US zip codes, using SOC codes, seniority, and company features, wired into an MLflow-tracked, auto-recalibrating ClickHouse pipeline for real-time forecasts.
  • Built PyOD (Isolation Forest, KNN) and Cleanlab validation over 15M+ distributed records, cutting false-positive anomaly flags by 35% in production.
  • Designed Polars / PySpark ingestion across ClickHouse, PostgreSQL, and AWS S3 to keep the job-market data clean for everything downstream.
LightGBMMLflowClickHousePyODCleanlabPolarsPySparkAWS S3
Mar 2023 – Jul 2024
AastraZen Technologies
Chennai, India

Data & Analytics Engineer

  • Fine-tuned DistilBERT to auto-categorize support tickets and added vector semantic search, dropping manual triage by 45%.
  • Migrated a legacy database to a hybrid MongoDB + PostgreSQL setup with dbt transformation layers; query throughput went up 40% and reporting went real-time.
  • Stood up a Kafka streaming pipeline into an AWS S3 data lake serving 5+ client accounts at 99.5% uptime.
DistilBERTMongoDBPostgreSQLdbtApache KafkaAWS S3
Toolkit

What I reach for

AI & Machine Learning

LLM Orchestration (Claude, MCP)Generative AILightGBMXGBoostSHAPIsolation ForestARIMASetFitBERTopicDistilBERTScikit-learn

Data Platforms & Warehousing

SnowflakeClickHousePostgreSQLDatabricksRedshiftMongoDBDuckDBAWS (S3, EC2, SAA-C03)MLflowdbtKafka

Languages & Processing

PythonSQLRPandasNumPyPolarsPySparkPyODCleanlab

Visualization & Reporting

TableauPower BIQlik SenseSigmaQuarto & JupyterKPI TrackingRoot Cause Analysis
Projects

Things I built to find out if I could

Agentic AIDec 2025

TravelGenie: AI travel planner

An agentic trip planner that orchestrates 7 MCP servers through Claude, pulling live flights, hotels, and events via SerpAPI into a single request. I load-tested EC2 against ECS Fargate at 50 to 100 concurrent users; Fargate matched EC2's 460 RPS on half the compute.

ClaudeMCPSerpAPIAWS ECS / EC2Load Testing
View on GitHub →
Computer VisionMar 2023

Farmland wildlife intrusion detection

A YOLO computer-vision pipeline that flags wildlife intrusions on farmland in real time and fires the alert automatically, so no one is stuck watching a monitor. It handles multi-class classification straight off live camera feeds.

★ 2nd of 500+ teams · National Technocraft Hackathon · published in BIOGECKO (Vol 12, Iss 03, 2023)
YOLOv5CNNPythonIoT
View on GitHub →
Time Series2025

Bitcoin price forecasting

End-to-end pipeline: live BTC data collection, ARIMA + XGBoost models, automated ETL, and interactive Qlik Sense dashboards. Landed around 85% accuracy and cut the manual work by 40%.

ARIMAXGBoostQlik SensePython
View on GitHub →
Analytics2025

Sage Ventures: multifamily analytics

Python and SQL analytics over a multifamily portfolio, quantifying $2.1M+ of loss-to-lease and a 26-point occupancy gap across properties, units, and rent rolls, with automated Excel and Power BI reporting for stakeholders.

PythonSQLPower BIExcel
View on GitHub →
Education

Where I learned it

Master's in Data Science

University of Maryland
May 2026 · Maryland
GPA 3.74 / 4.0
Natural Language Processing · Machine Learning · Big Data Systems · Deep Learning

B.Tech, Electrical & Electronics Engineering

Sri Krishna College of Engineering & Technology
Aug 2020 – May 2024 · Coimbatore, India
CGPA 8.4 / 10.0
Microcontrollers · Power System Analysis · Python · ML in Energy Systems
Let's build
something useful.

Open to full-time data science, ML engineering, and analytics roles. The inbox is always on.