Hi, my name is

Shubham Chawla

AI & Data Engineer, Google

I am an AI & Data Engineer at Google based in Hyderabad. I specialize in architecting distributed data lakehouses and production agentic AI systems for major banks and enterprise financial workloads.

Shubham Chawla

AI & Data Engineer @ Google

Hyderabad, India
Big Data Lakehouses & Multi-TB Scalable Pipelines
Agentic AI, GraphRAG & Model Context Protocol
Knowledge Graphs using Graph Databases (Spanner & Neo4j)
Agentic AI Platforms on Google Cloud (ADK)
Application Deployment & Architecture on GCP Infra
Production-Ready Agentic AI for Large Enterprises
Distributed Streaming & OpenTelemetry Tracing

Who am I?

I am an AI & Data Engineer at Google with over 10+ years of global experience architecting mission-critical data platforms and distributed systems.

My engineering focuses on building high-throughput Big Data lakehouses and operationalizing production-grade Agentic AI workflows—partnering with major banks and global financial institutions to deliver scalable, secure, and deterministic architectures.

Companies Worked With2016 – Present
GoogleCurrent
2022 – Present
Carelon
2021 – 2022
D
Deloitte
2018 – 2021
Cognizant
2016 – 2018

Featured Projects

A selection of production lakehouses, GraphRAG engines, and open Model Context Protocol tooling.

AGENTIC AIOpen Source

Equity Portfolio Analyzer Agent (Google ADK)

Multi-source financial agent built on Google ADK fanning out across 4 market data feeds (Yahoo Finance, Twelve Data, NSE/BSE, Tavily) to synthesize an opinionated BUY / HOLD / SELL verdict.

4 Parallel
Data Feeds
~115 Metrics
Data Points
Google ADK
Framework
Google ADKLiteLLMPythonFastAPIFinancial APIsNSE/BSE
MCP TOOLINGOpen Source

Production Database Model Context Protocol (MCP) Suite

Enterprise MCP servers enabling AI agents to safely introspect schemas, generate optimized SQL, and execute database tasks with strict guardrails.

< 140ms
Invocation Latency
100%
Safety Pass Rate
Cloud Run / Docker
Supported Runtimes
Model Context Protocol (MCP)FastAPIGoogle Cloud RunBigQueryCloud SQLDocker
DATA ENGINEERINGOpen Source

GCP Anti-Money Laundering (AML) AI Framework

Enterprise Anti-Money Laundering (AML) transaction compliance pipeline and dashboard deployed on GCP using FastAPI, BigQuery, and Cloud SQL.

Banking AML
Target Domain
BigQuery / SQL
Query Engine
FastAPI / GCP
Backend
Google CloudBigQueryFastAPICloud SQLPythonSQL DDL/DML
AGENTIC AIOpen Source

Agentic AI Observability & Telemetry Framework

Distributed tracing and evaluation framework tracking multi-agent reasoning steps, tool latencies, and token cost telemetry into GCP Cloud Trace.

< 4ms
Trace Overhead
Near Real-Time
Eval Pipeline SLA
Per-Agent Granular
Cost Attribution
OpenTelemetryGoogle Cloud TracePythonVertex AI PipelinesPrometheus
AGENTIC AIOpen Source

Automated Knowledge Graph Construction via LLM

Pipeline leveraging few-shot prompt chaining and graph reconciliation to convert unstructured domain documents into queryable property graphs.

91.8%
Entity Extraction F1
100k+ docs/day
Processing Speed
Strict Schema
Ontology Alignment
Gemini 1.5 ProPythonSpanner GraphPydanticNetworkXGQL
AGENTIC AIOpen Source

SQL Reverse-Engineering & Data Model Agent

AI agent that parses complex enterprise SQL queries, reverse-engineers entity-relationship models, and visualizes schemas in structured JSON.

Structured JSON
Output Format
Streamlit
UI Interface
FastAPI + AI
Architecture
PythonFastAPIStreamlitSQL ParserGoogle Cloud SDKPoetry
DATA ENGINEERINGProduction

High-Throughput CDC & Lakehouse Physical Modeling

Real-time Change Data Capture (CDC) streaming into BigQuery with partitioned physical modeling and LookML semantic metric layers.

50k+ eps
Throughput
38%
Query Cost Reduction
99.9%
Data Quality Score
BigQuery CDCPySparkDataplexLookMLCloud StoragePyTest
MCP TOOLINGProduction

Cloud Run Microservices & Semantic Agent API

Containerized, auto-scaling API services on Google Cloud Run delivering low-latency semantic search and analytical agent endpoints.

< 800ms
Cold Start
99.99%
Availability
Thousands req/min
Throughput
FastAPIGoogle Cloud RunDockerCloud SQLPythonGCP

Work Experience

A decade of engineering distributed data systems, enterprise lakehouses, and production Agentic AI across Google, Carelon, Deloitte, and Cognizant.

AI & Data Engineer

Google
Hyderabad, India•2022 – Present
Current Role • Featured Impact

Partnering with major banks, financial institutions, and global enterprises to architect mission-critical Big Data lakehouses and deploy production Agentic AI systems with strict security guardrails.

Key Engineering Deliverables & Impact at Google

Engineered production agentic AI workflows and Model Context Protocol (MCP) database toolboxes for financial workloads with AST-level query validation.
Designed sub-second GraphRAG retrieval pipelines combining Cloud Spanner Graph and BigQuery for complex multi-hop financial knowledge reasoning.
Implemented high-volume CDC streaming and lakehouse physical models processing 50k+ events/sec benchmarked against enterprise TPC-DS workloads.
Pioneered end-to-end LLM agent observability using OpenTelemetry to capture reasoning traces, latency, token economics, and evaluation scores.
Core Stack:BigQueryCloud SpannerPySparkLangGraphModel Context ProtocolVertex AIOpenTelemetryCloud Run

Previous Enterprise Experience

Senior Data Engineer

Carelon

Engineered scalable Big Data pipelines and cloud analytics infrastructure for massive healthcare and enterprise workloads.

2021 – 2022Hyderabad, India
Built distributed PySpark pipelines processing multi-TB transactional healthcare and claims records.
Engineered automated ETL workflows and validation frameworks ensuring zero data drift.
Optimized query runtimes and partitioning models on distributed storage for executive reporting.
Collaborated across international engineering teams on production cloud data platforms.
PySparkApache SparkPythonCloud PlatformsSQLAirflow

Articles & Writing

Deep dives published in Google Cloud on Medium covering Agentic AI, OpenTelemetry observability, and cloud data migrations.

Published on Medium

Explore more engineering articles & tutorials

Deep dives on Google Cloud Platform, production Agentic AI workflows, BigQuery, and enterprise lakehouse architectures.

Visit Medium Profile

Skills & Technologies

Core competencies across enterprise data pipelines, agentic workflows, and Google Cloud platforms.

Big Data & Distributed Systems

Enterprise storage engines, lakehouses, streaming ingestion, and scalable data transformations.

Google BigQueryAdvanced
Google Cloud SpannerAdvanced
PySpark & Spark SQLAdvanced
Google Cloud DataplexProficient
Cloud SQL & PostgreSQLProficient
Data ModelingAdvanced
Vertex AI PipelinesProficient
LookML & Semantic LayersProficient

Agentic AI & LLM Systems

Autonomous reasoning loops, multi-agent supervision, graph retrieval, and tool ecosystems.

Model Context Protocol (MCP)Advanced
GraphRAG & Knowledge GraphsAdvanced
Agent Development Kit (ADK)Advanced
LangGraph / Multi-Agent SupervisionAdvanced
Tool Calling & Function OrchestrationAdvanced
Vector Databases & Semantic RetrievalProficient
LLM EvaluationProficient
GQL (Graph Query Language)Proficient

GCP & Agentic Platform Expert

Enterprise Google Cloud architecture, specialized in the GCP Agentic Platform, Vertex AI, serverless runtimes, and production observability.

Google Cloud Platform (GCP)Expert
GCP Agentic PlatformExpert
Vertex AI Agent BuilderAdvanced
Google Cloud Run & ServerlessAdvanced
Cloud Architecture & SecurityAdvanced
OpenTelemetry & Cloud TraceAdvanced
FastAPI & MicroservicesAdvanced
Docker & ContainerizationProficient

Contact

Get in touch!

Whether you want to discuss distributed data architectures, agentic AI frameworks, or potential collaborations, my inbox is always open!

shubhamchawla10@gmail.com
Based in Hyderabad, India