Applied AI engineer · researcher

Akshat Gupta

I build machine-learning systems and spend most of my time finding out where they break.

Right now that means agentic AI and document intelligence for finance and insurance: systems that read messy, high-stakes documents and have to get the answer right, show their evidence, and survive an audit. My days move between reading papers, designing evaluations, and shipping production code.

Before LLMs took over, I worked across NLP, speech, computer vision, and adversarial ML. I studied Computational Linguistics at the University of Stuttgart and have been building with ML since 2018.

Akshat Gupta
Stuttgart, Germany

Three projects that show how I work: one production system, one published paper, one experiment that's still open.

AI systems for document-heavy workflows

Financial and insurance documents are long, inconsistent, and expensive to get wrong. I build the pipelines that read them — OCR, retrieval, specialised agents, validation — and return structured output a human reviewer can check. The interesting problems are rarely the model; they're grounding and knowing when to hand off to a person.

Work details →

  • Agentic AI
  • Azure OpenAI
  • OCR
  • Evaluation

GlyphNet

Phishing domains that look like trusted ones — paypa1.com with a digit for an L — can slip past text-based filters. GlyphNet renders the domain as an image and uses an attention-based CNN to spot the visual trick. Published with a 4M-domain dataset; 0.93 AUC.

Project website ↗Paper ↗Code ↗Dataset ↗

AUC
0.93
Original dataset
4M domains

SpeakerDiff

An open question I keep coming back to: can diffusion models generate speaker embeddings that anonymise a voice without wrecking it? SpeakerDiff is the prototype for testing that. It's not a finished answer — the privacy claim needs evaluation of the full conversion pipeline — but the code and experiments are public.

Code and experiments ↗

  • Diffusion models
  • Speech
  • Privacy

Two datasets I've released so others don't have to build them.

Working notes — things I'm learning, systems I'm building, and questions I haven't closed yet.

All writing →
AI Engineer · additiv

Agentic document and decision systems for finance and insurance — from OCR to production API.

Machine Learning Engineer · Validaitor

Testing models for the ways they fail: robustness, adversarial inputs, fairness, bias, and toxicity.

Research Assistant · University of Stuttgart

How language models represent time and place.

Machine Learning Engineer

NLP, machine translation, biomedical NER, speech, diarization, and multimodal learning.

Complete experience →

Away from the keyboard

Table tennis, chess, cycling around Stuttgart, following the markets, and reading more quantum physics than is strictly useful. There's usually a notebook involved.