luv. Let's Talk

Open to 2027 graduate roles

Hi, I'm Luv

An AI and data engineer building retrieval systems that show their sources.

Open graphite mechanical keyboard case with a screen in the lid and an orange hinge
LUV AGRAWAL / CREATIVE INDEXTYPING

Retrieval you can check. Pipelines that ship.

  • 2 internships
  • ✦
  • 94/100 GraphRAG accuracy at k=5
  • ✦
  • 100,500 records in under 2 seconds
  • ✦
  • 90%+ degree average
  • ✦
0query accuracy gain, GraphRAG over a hybrid baseline (54/100 to 94/100 at k=5)
0graph nodes and 57,050 edges across 1,000+ NHS documents
0synthetic patient records processed in under 2 seconds on AWS Lambda
0UK Tier 1 bank as first client of the contract intelligence platform

Experience

Two internships, shipped in parallel

HaloRFP · July to September 2026

From blank page to a bank as first client

AI and Data Engineering Intern. I built a contract intelligence platform for commercial real estate and construction, from initial design to first client deployment.

  • PDF and DOCX ingestion with clause-aware chunking of legal text
  • LLM extraction into a JSON schema of obligations, break clauses, payment terms and risk flags
  • Fine-tuned a small Qwen3-4B language model
  • Every answer cited back to its exact source passage
  • Full technical handover so the team could keep building
Ingestion · live pass1 / 3
PDFTender pack
ObligationsBreak clausesPayment termsRisk flags
Cited to source passage
I/O Atelier · June to September 2026

GraphRAG that beat the baseline by 74%

AI and Data Engineering Intern. I built a GraphRAG layer over a BM25 and vector Hybrid RAG baseline and evaluated it on 500 clinical queries across five top-k settings.

  • Knowledge graph of 23,576 nodes and 57,050 edges persisted in pgvector
  • Retrieval confidence up 27 points, 81% vs 54%
  • Removed three classes of false positives before scaling to 1,000+ documents
  • Event-driven AWS ingestion with least-privilege IAM and Neo4j loading
Evaluation · k=5500 queries
0graph nodes
0graph edges
0confidence
University of West London · 2024 to 2027

BSc Computer Science, 90%+ average

Final year. Strong theory and mathematics foundations, which is where the retrieval and evaluation instincts come from.

  • BMW Group UK Data Analytics Final Assessment Centre, December 2025
  • Mathematics Olympiad: Distinction (Grades 5 and 7)
  • Microsoft x EDI Diversity in Tech Competition, 2025
  • AWS Cloud Practitioner Foundations: in progress
Key resultsFinal year

The graph

A knowledge graph you can walk

Dense knowledge graph visualisation: thousands of green nodes joined by grey edges, with a few large dark hub nodes at the bottom right
Knowledge graph visualisation. Each dot is an entity pulled from the documents, each line a relationship.

Plain vector search finds passages that look similar. It misses answers that live across several documents. The graph connects the entities those documents share, so a query can walk from one document to the next.

  • Entities extracted with DeepSeek-V3 and stored in pgvector
  • NetworkX traversal across 8 to 14 documents per query
  • Scored on 500 clinical queries across five top-k settings
  • Exposed to AI assistants as MCP tools, with no direct database access

How I build

From raw documents to answers you can check

01 Ingest

Parse everything, chunk by meaning

Contracts, leases and tender packs arrive as PDF and DOCX. I parse them, split legal text clause by clause rather than by character count, and keep parties, key dates and document type as metadata.

02 Extract

Turn text into a schema

An LLM converts clauses into a defined JSON schema: obligations, break clauses, payment terms and risk flags, ready for downstream legal review.

03 Retrieve

Hybrid search plus a graph

BM25 and vector search find candidates, then a knowledge graph traverses across 8 to 14 documents per query. Keyword-overlap thresholds, a confidence floor and citation recalibration cut the false positives.

04 Prove and hand over

Measure it, then document it

Every change is scored on hundreds of real queries. When the work ships, the team gets the architecture, runbook, troubleshooting guide, API reference, decisions log and cost model.

Ingestion · clause chunksmetadata
Clause
Clause
Clause
PartiesKey datesDocument type
Extraction schemaexcerpt
{
  "obligations": [ ... ],
  "break_clauses": [ ... ],
  "payment_terms": { ... },
  "risk_flags": [ ... ],
  "source_passage": "exact text"
}
Retrieval · graph traversal8 to 14 docs
BM25VectorGraph
Evaluation · k=5500 queries
ArchitectureRunbookTroubleshootingAPI referenceDecisions logCost model

Toolkit

What I build with

AI and RAG

  • GraphRAG
  • Hybrid RAG
  • BM25
  • Vector search
  • Knowledge graphs
  • NetworkX
  • Entity extraction
  • RAG evaluation
  • LLM fine-tuning
  • Qwen3
  • DeepSeek-V3
  • MCP
  • LangChain
  • Prompt engineering

Programming

  • Python
  • SQL
  • FastAPI
  • REST APIs
  • Pandas
  • NumPy
  • Boto3
  • GeoPandas
  • Matplotlib
  • Seaborn

Cloud

  • AWS
  • Lambda
  • S3
  • IAM
  • CloudWatch
  • EC2
  • Terraform
  • Serverless
  • Event-driven pipelines
  • Docker

Data and tools

  • PostgreSQL
  • Database sharding
  • Supabase
  • pgvector
  • Neo4j
  • Git and GitHub
  • Jupyter
  • ClickUp
  • Agile Scrum
  • Technical documentation

Good systems start with hello.

Have a role, a project in mind, or just want to say hi? I'm looking for a 2027 graduate role in AI or Data Engineering.

luvga07@gmail.com

GitHubLinkedInDownload CV

© 2026 Luv Agrawal · London, UK · Back to top