About me
Hands-on IT leader with 15+ years specializing in Big Data, Hadoop, Spark, Generative AI, LLMs, and cloud computing across Azure, AWS, and GCP. Expert in building scalable ETL pipelines, real-time streaming with Kafka, and integrating GenAI capabilities using LangChain/LangGraph with enterprise data platforms. Proven track record in Databricks, Snowflake, data quality frameworks, and production support. Skilled in PySpark, Scala, SQL, and modern orchestration tools. Strong problem-solver with deep expertise in SDLC Agile methodologies and cross-functional collaboration.
- 15+ years of IT experience specializing in Big Data, Hadoop, Spark, Gen AI, LLM, and multi-cloud computing
- Hands-on experience building LLM applications and AI agents using LangChain/LangGraph with enterprise data integration
- Integrated GenAI capabilities with Databricks, Spark, SQL, and cloud data pipelines for scalable, secure LLM solutions
- Expert in real-time data processing with Kafka and Spark Streaming for high-volume enterprise workloads
- Proven track record optimizing Spark jobs and implementing data quality frameworks using AWS Deequ
- Strong production support experience with fine-tuning, deployment, and troubleshooting in enterprise environments
Where I've worked
Technical Lead · Tech Mahindra
June 2023 – Oct 2025- Developed PySpark ETL pipelines processing high-volume product, sales, inventory, and customer datasets using Medallion Architecture (Bronze/Silver/Gold) in Databricks
- Built ingestion pipelines with StreamSets for CSV, JSON, relational databases, and APIs; implemented data quality checks and incremental processing using watermarking and CDC
- Orchestrated workflows with Airflow DAGs; loaded curated datasets into Snowflake with optimized schemas, clustering, pruning, and caching strategies
- Designed multi-cloud data pipelines using Azure Data Factory, AWS Glue, and SSIS; built scalable AWS ETL workflows with S3, Glue, Lambda, Athena, and Redshift
- Developed event-driven ingestion pipelines using SQS, Lambda, Step Functions, and Kafka-based real-time streaming
- Enforced data quality using AWS Deequ; optimized Spark jobs via partitioning, broadcast joins, and Adaptive Query Execution
Lead System Analyst · UST Global
September 2020 – May 2023- Migrated data from on-premise Cassandra and Teradata to Azure Data Lake Storage for Dell customer payments analytics
- Created ADF pipelines using Linked Services, Datasets, and copy activities to extract, transform, and load data from Azure SQL, Blob Storage, and SQL Data Warehouse
- Developed JSON scripts for deploying ADF pipelines; wrote SQL queries on Azure notebooks to analyze data stored in Azure Delta Lake
- Monitored and troubleshot Databricks clusters; created dashboards and visualizations based on insights from Azure Delta Lake
- Implemented optimization techniques and best practices in Azure Data Factory; created data flows for business rule transformations
- Deployed Python packages as wheel files in Databricks; scheduled jobs to extract and load data into ADLS Gen2 hourly
Senior Consultant · Capgemini
May 2017 – August 2020- Implemented backend workflow using Hadoop and Spark for PSA Automobiles' Product Lifecycle Management solution
- Developed Spark and Kafka jobs to receive data from Insurance SDK Chips; created Kafka topics based on source-generated events
- Created Spark-Cassandra and Spark-HDFS connectors to store structured data; created Hive tables to retrieve data from HDFS
- Performed functional and performance testing using MQTT Lens and JMeter; connected Hive to Tableau for dashboards and visualizations
- Designed and created infrastructure for Thrivent financial firm's credit collection platform migration from on-premise Hortonworks to HDFS and Cassandra
- Developed optimized Spark jobs to process data into understandable formats; stored data in HDFS and Cassandra
Senior Software Engineer · CenturyLink
May 2010 – April 2017- Developed VDSL Data Mart application integrating Choice TV, Choice Online subscriber, and operational data from various systems
- Supported Capacity Provisioning, Marketing, Finance, and Network Service Assurance teams with web reports for analytics and planning
- Managed data for VDSL systems in Denver and Phoenix, analog cable in Omaha, and Direct TV VMDU project across 14-state region
- Maintained data related to CSG billing system, WFA video, help desk reporting, loop qualification, capacity monitoring, and marketable homes
What I know
- Hadoop
- Spark
- PySpark
- Scala
- Spark SQL
- Kafka
- HDFS
- Hive
- Generative AI
- LLM
- LangChain
- LangGraph
- RAG
- Prompt Engineering
- Vector Databases
- Azure Data Factory
- Azure Databricks
- Azure Data Lake
- Azure Synapse
- AWS S3
- AWS Lambda
- AWS Glue
- AWS Bedrock
- AWS Redshift
- AWS EMR
- GCP BigQuery
- GCP Dataflow
- GCP DataProc
- Snowflake
- Cassandra
- MongoDB
- MySQL
- Oracle
- SQL Server
- Airflow
- Unity Catalog
- Power BI
- StreamSets
- Terraform
- CloudFormation
- Datadog
- Git
- JIRA
- Agile Scrum
- Shell Scripting
- JSON
- Delta
- ORC
- Parquet
- Avro
- CSV
Where I studied
B.Tech
Dhaneswar Rath Institute of Engineering & Management Studies (D.R.I.E.M.S), Orissa · 2006
Consulting Services
- Custom consulting packages tailored to client-specific needs
- Focus on building deep partnerships with measurable impact over time
- Rates vary by project type with detailed breakdowns provided after scope discussion
- Open to new projects with collaborative, tailored approach
Himansu's CV answers questions directly, from their own documents.