Data Engineer | Data Analyst

Bhargav Kodaru

Data Engineer and Marketing Data Analyst with 4+ years of experience designing SQL-driven audience segmentation, campaign data pipelines, marketing automation QA workflows, and analytics-ready datasets across Salesforce Marketing Cloud, HubSpot, Marketo, AWS, Azure, Snowflake, Databricks, dbt, Airflow, SQL, and Python environments.

  • 4+ years across marketing data engineering and analytics delivery
  • SQL + QA audience segmentation, suppression logic, and deployment validation
  • SFMC campaign-ready data flows, personalization fields, and lifecycle reporting
Portrait of Bhargav Kodaru in a navy suit
Currently focused on campaign data engineering, segmentation QA, governed ETL, and trusted reporting for marketing teams.

Marketing data engineering with practical delivery discipline

I specialize in translating lifecycle marketing requirements into reliable SQL queries, reusable segmentation filters, suppression logic, personalization attributes, contact eligibility rules, and governed campaign-ready data extracts.

My background combines ETL automation, code governance, peer review, campaign validation, data integrity checks, and cloud data warehousing so teams can deploy with more confidence and trust their downstream reporting.

High-impact strengths for campaign data delivery

Campaign Logic

Segmentation and targeting

  • SQL audience segmentation
  • Email campaign logic
  • APEX campaign validation
  • Audience suppression and overlap checks
  • Contact eligibility logic

QA & Governance

Reliable campaign deployment

  • Campaign QA frameworks
  • Code governance and peer review
  • Version control and branching workflows
  • Data integrity validation and reconciliation
  • Deployment readiness checks

Core stack

Languages

SQL, Python, PySpark, R, Scala

Marketing Platforms

Salesforce Marketing Cloud, HubSpot, Marketo, Salesforce CRM, Journey Builder, Automation Studio

Campaign Data Engineering

Audience Segmentation, Email Campaign Logic, APEX Campaign Validation, Personalization Logic, Suppression Rules, Contact Eligibility Checks

Cloud Platforms

AWS, Azure, GCP

AWS Services

AWS Glue, Amazon S3, Amazon Athena, AWS Lambda, EventBridge, Amazon Redshift, RDS, Glue Data Catalog, IAM, AWS Step Functions

Data Warehousing & Databases

Snowflake, Amazon Redshift, Azure Synapse, BigQuery, SQL Server, PostgreSQL, MySQL, Azure SQL Database

Analytics Engineering

dbt Models, Dimensional Modeling, Star Schema, Fact Tables, Dimension Tables, KPI Datasets, Reporting Datasets

Orchestration & Workflow

Apache Airflow, Azure Data Factory, AWS Step Functions, EventBridge, Informatica/IICS, StreamSets

QA, Governance & Review

Code Governance, Peer Review, Pull Requests, Git Version Control, Schema Validation, Duplicate Detection, Null Checks, Reconciliation, Audit Logs

Monitoring & Reliability

SLA Monitoring, Error Logging, Retry Logic, Failure Notifications, Job Status Tracking, Root Cause Analysis, Idempotent Loads

Reporting & BI

Power BI, Tableau, DOMO, DAX, Power Query, Excel, Dashboard Models, Campaign Performance Reporting

Experience aligned to marketing data and campaign QA

Jan 2025 - Present USA

Data Analytics Engineer

Wesco International

  • Designed and maintained scalable SQL-driven audience segmentation and analytics datasets using Azure Data Factory, Azure Databricks, PySpark, Snowflake, and dbt, enabling lifecycle marketing, campaign reporting, personalization, and operational analytics across high-volume customer and transaction data.
  • Built reusable email campaign logic, segmentation filters, contact eligibility rules, suppression logic, personalization attributes, and campaign-ready extracts for Salesforce Marketing Cloud workflows, improving targeted audience accuracy by 38% across recurring marketing initiatives.
  • Developed validation workflows for APEX campaigns and advanced campaign structures by testing audience counts, SQL joins, exclusion logic, channel suppression rules, overlap conditions, and personalization fields before deployment, reducing campaign logic defects by 32%.
  • Implemented automated campaign QA checks using Python, SQL, dbt, and Great Expectations to validate schema consistency, duplicate contacts, null personalization values, invalid email attributes, audience mismatches, and source-to-target reconciliation across campaign datasets.
  • Established code governance workflows using Git, pull requests, SQL review templates, branching standards, reusable query patterns, and peer-review checkpoints, improving version control and reducing manual campaign logic changes across production deployment cycles.
  • Conducted technical code reviews for segmentation queries, dbt models, campaign filters, and audience extracts to verify logic accuracy, data integrity, business-rule alignment, and audience consistency before email and APEX campaign execution.
  • Troubleshot data discrepancies, audience overlap issues, logic errors, and campaign count mismatches by analyzing SQL joins, source-to-target mappings, incremental load dependencies, refresh schedules, aggregation rules, and marketing automation data outputs.
  • Optimized Snowflake, Azure Synapse, and Redshift workloads through partitioning, clustering, SQL tuning, dimensional modeling, and curated campaign data marts, improving query performance by 40% for BI dashboards, campaign analysis, and operational reporting.
Jan 2022 - Dec 2023 India

Data Engineer

Persistent Systems

  • Designed and maintained AWS-based ETL pipelines using AWS Glue, Amazon S3, Amazon Athena, SQL, PySpark, and Redshift to ingest, transform, and publish curated datasets for marketing analytics, campaign reporting, and business intelligence consumption.
  • Built incremental ingestion pipelines from MySQL RDS, APIs, flat files, CRM exports, marketing automation data, and cloud storage into S3 lakehouse zones, reducing redundant processing by 35% and improving refresh efficiency for scheduled campaign and analytics workflows.
  • Developed reusable PySpark and SQL transformation logic to standardize source data, handle schema changes, apply segmentation rules, and generate curated Parquet datasets optimized for Athena, Redshift, Tableau, Power BI, and downstream campaign reporting tools.
  • Engineered audience-ready data models with customer attributes, engagement signals, transaction metrics, campaign history, eligibility indicators, suppression flags, and KPI fields to support lifecycle marketing segmentation, personalization analysis, and targeted reporting use cases.
  • Implemented reconciliation and data integrity checks using duplicate detection, schema validation, null handling, SQL lookups, record-count balancing, and invalid contact checks, maintaining 99% accuracy across critical ETL workflows and reporting datasets.
  • Integrated AWS Glue Data Catalog for centralized metadata management, schema discovery, governed table access, and reusable dataset documentation, improving consistency and traceability across analytics, segmentation, and campaign data pipelines.
  • Automated serverless workflow execution using AWS Lambda, EventBridge, and AWS Step Functions to trigger Glue jobs dynamically, automate campaign data refresh cycles, and reduce manual intervention in recurring production ETL execution.
  • Optimized Athena and Redshift workloads by converting raw data into partitioned Parquet formats, tuning SQL queries, improving table design, and reducing query latency and cloud processing costs for downstream campaign analytics and BI users.
Jun 2021 - Dec 2021 India

Junior Data Engineer

TCS

  • Supported a marketing technology client by developing Azure Data Factory and Azure Databricks pipelines to ingest customer, campaign, engagement, and operational datasets into ADLS Gen2 for segmentation, reporting, and lifecycle marketing analytics use cases.
  • Built SQL and Spark-based validation logic to verify completeness, accuracy, consistency, schema conformity, duplicate conditions, and invalid audience records before loading curated datasets into downstream reporting and campaign execution workflows.
  • Developed reusable ingestion and transformation logic for files from AWS S3, Azure SQL Database, and ADLS Gen2, improving scalability and consistency across recurring customer, campaign, and analytics data pipeline executions.
  • Used Azure SQL Database for lookup validation, reference data checks, segmentation rule support, contact eligibility checks, and transformation logic, improving reliability of multi-source data integration across batch pipeline workflows.
  • Applied secure authentication and access controls using Azure Key Vault, managed identities, and role-based permissions to protect sensitive customer, campaign, and operational datasets across cloud data pipelines.
  • Supported code governance, QA documentation, and deployment validation by maintaining reusable SQL scripts, validating source-to-target mappings, checking audience counts, documenting campaign logic, and supporting peer review for campaign-ready datasets.

Selected work tied to reporting, automation, and audience insight

Project

Data Pipeline Automation with Apache Airflow

Designed and implemented an Airflow-based ETL workflow for scalable ingestion, validation, warehouse loading, and reporting readiness.

  • Ingested JSON data from Amazon S3 into Amazon Redshift using automated Airflow pipelines.
  • Developed custom DAGs and operators for staging, transformation, validation, and warehouse loading.
  • Improved pipeline reliability, repeatability, and data quality for recurring analytics processing and dashboard delivery.
Apache Airflow Amazon S3 Amazon Redshift ETL

Project

Predicting Heart Disease Risk Using Machine Learning

Built and tuned machine learning models to assess heart disease risk from patient health indicators using cleaned, validated, model-ready data.

  • Developed Logistic Regression and SVC models for classification.
  • Improved model performance with feature engineering, data cleaning, exploratory data analysis, and preparation of high-quality training datasets.
  • Applied undersampling techniques to address class imbalance and improve model quality.
Python Logistic Regression SVC EDA

Research-backed analytics interest

Predicting and Enhancing the Fluctuations of Cryptocurrency

Bhargav Kodaru

Research work focused on predictive modeling, analytical interpretation, and algorithm-driven approaches to cryptocurrency market fluctuations.

View Publication

Cloud and pipeline credentials

Microsoft

Azure Fundamentals (AZ-900)

Foundational Microsoft Azure certification covering core cloud concepts, Azure services, governance, pricing, security, and compliance basics.

View Credential

AWS

AWS Project Pipeline Certificate

TrendyTech certificate recognizing practical work in AWS pipeline design, cloud ETL concepts, and big data engineering foundations. Issued March 28, 2025.

StreamSets

StreamSets White Belt

Foundational certification covering data architecture basics, data ingestion, connectors, ETL concepts, pipeline execution, and platform operations.

StreamSets

StreamSets Yellow Belt

Hands-on certification focused on building pipeline solutions, using connectors, executing ETL workflows, and applying data integration concepts in delivery scenarios.

Academic background

Clark University

Worcester, Massachusetts

Master of Science in Data Analytics

Jan 2024 - Dec 2025

Relevant coursework: Mathematical Statistics, Regression Analysis, Advanced Programming in Python and R, Modern Data Engineering, Applied Machine Learning, Data Mining, Database Management, Data Visualization, Predictive Analytics, and Business Intelligence.

Let's build with data

If you are hiring for marketing data engineering, campaign QA, SQL segmentation, ETL automation, or analytics-focused data roles, I'd be glad to connect.