Himansu Sandha

Cloud Data Engineer | 15+ Years in Big Data, Spark, Hadoop, Gen AI & Multi-Cloud Solutions

Himansu Sandha

documents read
2
documents read
passages indexed
19
passages indexed
questions answered
11
questions answered
builder, since Sep 2026
#12
builder, since Sep 2026

About me

Hands-on IT leader with 15+ years specializing in Big Data, Hadoop, Spark, Generative AI, LLMs, and cloud computing across Azure, AWS, and GCP. Expert in building scalable ETL pipelines, real-time streaming with Kafka, and integrating GenAI capabilities using LangChain/LangGraph with enterprise data platforms. Proven track record in Databricks, Snowflake, data quality frameworks, and production support. Skilled in PySpark, Scala, SQL, and modern orchestration tools. Strong problem-solver with deep expertise in SDLC Agile methodologies and cross-functional collaboration.

  • 15+ years of IT experience specializing in Big Data, Hadoop, Spark, Gen AI, LLM, and multi-cloud computing
  • Hands-on experience building LLM applications and AI agents using LangChain/LangGraph with enterprise data integration
  • Integrated GenAI capabilities with Databricks, Spark, SQL, and cloud data pipelines for scalable, secure LLM solutions
  • Expert in real-time data processing with Kafka and Spark Streaming for high-volume enterprise workloads
  • Proven track record optimizing Spark jobs and implementing data quality frameworks using AWS Deequ
  • Strong production support experience with fine-tuning, deployment, and troubleshooting in enterprise environments

Where I've worked

  1. Technical Lead · Tech Mahindra

    June 2023 – Oct 2025
    • Developed PySpark ETL pipelines processing high-volume product, sales, inventory, and customer datasets using Medallion Architecture (Bronze/Silver/Gold) in Databricks
    • Built ingestion pipelines with StreamSets for CSV, JSON, relational databases, and APIs; implemented data quality checks and incremental processing using watermarking and CDC
    • Orchestrated workflows with Airflow DAGs; loaded curated datasets into Snowflake with optimized schemas, clustering, pruning, and caching strategies
    • Designed multi-cloud data pipelines using Azure Data Factory, AWS Glue, and SSIS; built scalable AWS ETL workflows with S3, Glue, Lambda, Athena, and Redshift
    • Developed event-driven ingestion pipelines using SQS, Lambda, Step Functions, and Kafka-based real-time streaming
    • Enforced data quality using AWS Deequ; optimized Spark jobs via partitioning, broadcast joins, and Adaptive Query Execution
  2. Lead System Analyst · UST Global

    September 2020 – May 2023
    • Migrated data from on-premise Cassandra and Teradata to Azure Data Lake Storage for Dell customer payments analytics
    • Created ADF pipelines using Linked Services, Datasets, and copy activities to extract, transform, and load data from Azure SQL, Blob Storage, and SQL Data Warehouse
    • Developed JSON scripts for deploying ADF pipelines; wrote SQL queries on Azure notebooks to analyze data stored in Azure Delta Lake
    • Monitored and troubleshot Databricks clusters; created dashboards and visualizations based on insights from Azure Delta Lake
    • Implemented optimization techniques and best practices in Azure Data Factory; created data flows for business rule transformations
    • Deployed Python packages as wheel files in Databricks; scheduled jobs to extract and load data into ADLS Gen2 hourly
  3. Senior Consultant · Capgemini

    May 2017 – August 2020
    • Implemented backend workflow using Hadoop and Spark for PSA Automobiles' Product Lifecycle Management solution
    • Developed Spark and Kafka jobs to receive data from Insurance SDK Chips; created Kafka topics based on source-generated events
    • Created Spark-Cassandra and Spark-HDFS connectors to store structured data; created Hive tables to retrieve data from HDFS
    • Performed functional and performance testing using MQTT Lens and JMeter; connected Hive to Tableau for dashboards and visualizations
    • Designed and created infrastructure for Thrivent financial firm's credit collection platform migration from on-premise Hortonworks to HDFS and Cassandra
    • Developed optimized Spark jobs to process data into understandable formats; stored data in HDFS and Cassandra
  4. Senior Software Engineer · CenturyLink

    May 2010 – April 2017
    • Developed VDSL Data Mart application integrating Choice TV, Choice Online subscriber, and operational data from various systems
    • Supported Capacity Provisioning, Marketing, Finance, and Network Service Assurance teams with web reports for analytics and planning
    • Managed data for VDSL systems in Denver and Phoenix, analog cable in Omaha, and Direct TV VMDU project across 14-state region
    • Maintained data related to CSG billing system, WFA video, help desk reporting, loop qualification, capacity monitoring, and marketable homes

What I know

  • Hadoop
  • Spark
  • PySpark
  • Scala
  • Spark SQL
  • Kafka
  • HDFS
  • Hive
  • Generative AI
  • LLM
  • LangChain
  • LangGraph
  • RAG
  • Prompt Engineering
  • Vector Databases
  • Azure Data Factory
  • Azure Databricks
  • Azure Data Lake
  • Azure Synapse
  • AWS S3
  • AWS Lambda
  • AWS Glue
  • AWS Bedrock
  • AWS Redshift
  • AWS EMR
  • GCP BigQuery
  • GCP Dataflow
  • GCP DataProc
  • Snowflake
  • Cassandra
  • MongoDB
  • MySQL
  • Oracle
  • SQL Server
  • Airflow
  • Unity Catalog
  • Power BI
  • StreamSets
  • Terraform
  • CloudFormation
  • Datadog
  • Git
  • JIRA
  • Agile Scrum
  • Shell Scripting
  • JSON
  • Delta
  • ORC
  • Parquet
  • Avro
  • CSV

Where I studied

  • B.Tech

    Dhaneswar Rath Institute of Engineering & Management Studies (D.R.I.E.M.S), Orissa · 2006

Consulting Services

  • Custom consulting packages tailored to client-specific needs
  • Focus on building deep partnerships with measurable impact over time
  • Rates vary by project type with detailed breakdowns provided after scope discussion
  • Open to new projects with collaborative, tailored approach
DiscoverableFindable on Google & AI search
One step ahead

Himansu's CV answers questions directly, from their own documents.

Builder #12Building since Sep 2026