Aman Jain

SWE III – AI/ML engineer at Google. I build ranking, personalization and LLM systems that move real metrics.

I'm based in Hyderabad. Before Google I worked on post-training for GPT-4o mini at OpenAI, rebuilt ranking and personalization at Cars24, built an agentic trading platform as founding AI engineer at KotiLabs, and ran entity resolution on a 900M-node graph at Vieu. I started as a data scientist at 6sense.

What I care about: systems that move a metric you can check, and designs where the model can't do anything unsafe even when it's wrong.

Selected work

Each has a deep-dive on the blog
+150% Click Recall@50, 12% → 30%

Per-user ranking that replaced 25 cohorts

Cars24 · Data Scientist · 2024–2025

Problem
Every user in a cohort saw the same ranking of ~10,000 cars, and a third of users had no clicks to personalise from.
What I did
Rebuilt the recommender end to end: two-tower retrieval with an HNSW index, a GBDT ranker, and cold-start from search and filter signals. Shipped behind a live A/B test.
Result
+150% Click Recall@50, +16% buyer conversion, personalised coverage from 65% to 100% of users, under 100ms p99. Later adopted in Australia, Thailand and India.
Read the deep-dive →
$0.10 → $0.01 LLM cost per query

An agentic trading platform built for SEBI compliance

KotiLabs (āagman) · Founding AI Engineer · 2025–2026

Problem
Let users trade by voice or chat without any chance that a language model places a non-compliant order.
What I did
Designed a 5-layer system where LLMs only plan: plans compile to a JSON DSL, and an OPA/Rego policy layer and a deterministic engine decide what runs. Added a semantic cache and a small intent router in front of the LLM.
Result
Zero hallucination-induced compliance breaches by design, over 90% of frontier-model calls avoided, and a screener running 500+ instruments at more than 10× sequential throughput.
Read the deep-dive →
900M person nodes in the graph

Entity resolution on a live relationship graph

Vieu · Data / ML Engineer · 2026

Problem
Link names scraped from athletics rosters and research papers to the right person in a 900M-row production table, where a wrong match invents a relationship.
What I did
Built precision-first canonicalization: ground-truth checks first, tiered candidate filtering to protect the live database, then scoring on name, school, year, company and hometown.
Result
Shipped the co-athlete edge type and designed the co-author pipeline over 250M+ academic records, with thresholds tuned to under-resolve rather than mis-resolve.
Read the deep-dive →

Also

  • LLM post-training, OpenAI. Owned the post-training stack for GPT-4o mini, optimized inference latency and cost across model families, and helped define deployment criteria for new model releases.
  • Car price prediction, Cars24. About 70% of predictions within ±5% of the final sale price, using LLM-structured inspection data.
  • Default sort ranking, Cars24. +21% view-to-buy-intent on the top five deciles, confirmed in a live A/B test.
  • Logo similarity service, 6sense. 400M+ comparisons per run for Fortune 500 data deliveries, with 4% fewer false positives.

Career

Oct 2026 – now
Google · Software Engineer III, AI/MLBuilding AI/ML systems. Hyderabad.
Apr – Oct 2026
Vieu · Data / ML EngineerEntity resolution and graph pipelines for a 900M-node person graph. Remote.
  • Built the academic co-authorship edge pipeline (pubnet), a new edge type alongside co-worker, co-student and co-athlete, ingesting 250M+ entities from OpenAlex and OAG with deduplication and versioned edge scoring.
  • Designed a multi-stage canonicalization framework that matches scraped entities to ~900M person nodes in Postgres, with tiered candidate filtering to protect the production database under live traffic.
  • Ran data-source feasibility studies across OpenAlex and OAG, including canon-rate reconciliation and a 74% discard-rate analysis that shaped the pipeline design.
  • Worked across a pre-computed graph architecture (Postgres + OpenSearch) with Lambda ingestion, event-driven edge computation and path-finding for sales outreach.
Jul 2025 – Apr 2026
KotiLabs · Founding AI Engineer, āagman0→1 agentic trading platform under SEBI's algo-trading rules. Pre-seed, 10 people. Bengaluru.
  • Architected a 5-layer neuro-symbolic platform: LLM planning (Mastra AI) with deterministic OPA/Rego policy enforcement, so no hallucination can cause a regulatory breach.
  • Designed a JSON DSL as the LLM's only output format: the model picks building blocks and an interpreter runs them, removing arbitrary code execution and prompt-injection risk.
  • Built a VectorBT backtesting engine with golden test suites that check exact trade counts, timestamps and metrics.
  • Built a vectorized screener for 500+ instruments at more than 10× sequential throughput, with ClickHouse + TimescaleDB ingestion and adapters for Zerodha, Groww, Angel One and Upstox.
  • Built hybrid BM25 + dense retrieval memory (pgvector, HNSW) and an intent router (distilled Llama-3 8B + Redis semantic cache) that avoids over 90% of full LLM calls, cutting cost per query from $0.10 to $0.01.
  • Mentored 2 engineers on deterministic execution design, policy authoring and golden-test methodology.
Feb 2024 – Feb 2026
OpenAI · IC, LLM ArchitectPost-training, inference cost and deployment criteria.
  • Owned the post-training stack for GPT-4o mini.
  • Optimized inference latency and cost across model families.
  • Worked with safety, policy and product teams to define deployment criteria for new model releases.
Dec 2023 – Jul 2025
Cars24 · Data Scientist, Ranking, Personalization & PricingRecommendation, ranking and pricing models for 1M+ monthly sessions. Gurugram.
  • User-level personalization that replaced a 25-cohort system: +150% Click Recall@50 (12% → 30%), +16% buyer conversion, under 100ms p99, and cold-start solved for the 35% of users with no clicks.
  • Default Sort V2 (GBDT) with dynamic de-boosting of underperforming inventory: +21% view-to-buy-intent on the top five deciles and +7% clicks per vehicle, validated out-of-time and in a live A/B test.
  • Final selling price model combining vehicle fingerprint, demand and supply, and LLM-structured inspection data: about 70% of predictions within ±5% of the actual sale price.
  • Hybrid "Similar Cars" recommendations (70:30 collaborative:content): +6% impressions and +12% buyer intent from that widget.
  • Migrated 9 models to GA4 with zero downtime and owned data science for the UAE, Thailand and Australia at the same time.
  • Mentored 2 junior analysts on A/B test validation and out-of-time evaluation.
May 2021 – Oct 2023
6sense · Data Scientist, intern → full-timeB2B intent data, NLP and computer vision. Bengaluru.
  • Logo-similarity microservice (embedding-based computer vision on AWS S3 + Hive): 400M+ comparisons per run, SLA-bound fortnightly deliveries to Fortune 500 clients, 4% fewer false positives.
  • Maintainer of Togylop, 6sense's internal NLP training library (BERT/RoBERTa classification and token tagging).
  • Scaled the B2B intent taxonomy from 57 to 198 divisions with supervised entity classification.
Dec 2021 – May 2023
Ministry of Science & Technology, Govt. of India · Data AnalystETL and policy-monitoring dashboards. New Delhi.
  • Built automated ETL pipelines and analytics dashboards for policy monitoring, consolidating departmental datasets into standard reports for senior officials.

Research

Comparative Study of BERT and Legal-BERT for Predicting Indian Legal Case Judgements. First author, from my M.Tech thesis. JURISIN 2022 workshop, published in Springer LNAI (2025). It shows that domain-specific pre-training clearly beats general BERT on Indian case-judgement prediction.

What I work with

ML & ranking
Ranking, personalization, LightGBM/GBDT, two-tower retrieval
LLMs & agents
Multi-agent orchestration, hybrid RAG, LoRA/PEFT fine-tuning, semantic caching
Data & infra
PostgreSQL/pgvector, ClickHouse, OpenSearch, Snowflake, Redis
Experimentation
User-level A/B tests, power analysis, Bayesian testing, SRM checks, out-of-time validation