Skip to content
Anton Dergunov

Projects

My own projects, grouped by what they are about. Each says where it stands, and work that is not done yet is listed as what comes next.

Evaluation and post-training

2 projects

LLM agent evaluation test bed

active

A support agent for a fictional neobank, built to iterate on how LLM agents are evaluated. The model only talks; plain Python makes every decision that touches money or account state, so prompt injection has no action to reach. An LLM plays the customer, and every simulated conversation is scored by deterministic checks and an LLM-as-judge.

Next The evaluation experiments written up as a six-part series.

  • llm evaluation
  • llm as judge
  • agent evaluation
  • ai agents
  • guardrails
  • rag
  • prompt injection
  • python

Mandarin Speech Coach

prototype

A different way to practise Mandarin tones. Where pronunciation assessment usually returns a score, this shows the learner's own pitch contour against the target phrase, aligned character by character and drawn in tone colours, so what went wrong is visible. A recording is transcribed, force-aligned and pitch-tracked to produce it.

Next An alignment evaluation harness, a surface-tone classifier, and a compact aligner fine-tuned to pinyin initials and finals.

  • speech
  • pronunciation
  • mandarin
  • forced alignment
  • wav2vec2
  • whisper
  • pitch tracking
  • language learning

Retrieval and ranking

3 projects

Spoken Usage Retrieval

active

An information-retrieval system over native speech: it finds real uses of a word or phrase and plays the exact moment each is spoken. An indexing pipeline builds the corpus from the captions of curated YouTube channels; retrieval matches words and short phrases by surface form and by lemma, with analysers compared across ten languages; results are ranked and diversified across videos. Experiments on aligning speech to subtitles and on word-level translation alignment are kept in the repo.

Next Human relevance judgments with a calibrated LLM judge, then multi-stage ranking: ranking features, a learned ranker and a multilingual reranker.

  • information retrieval
  • ranking
  • search
  • speech
  • forced alignment
  • nlp
  • language learning
  • python

Task Concept Retrieval

active

A retrieval test bed: given a task headline from a planner, return the one icon that conveys it, or nothing. The matching signal is generated: a vision model describes each of the 4,257 icons (what it shows, which tasks it suits, which it does not), and that description, never the icon's name, is what gets matched. No LLM runs on the query path, and a calibrated gate abstains when no icon is good enough.

Next A second generation of icon descriptions, a labelled gold set of real task headlines, and a matcher fine-tuned on them, benchmarked against embedding and reranking baselines.

  • information retrieval
  • semantic search
  • embeddings
  • sentence embeddings
  • reranking
  • calibration
  • python

Interest-aligned vocabulary recommendation

planned

Recommenders for language learning usually ask what a learner does not know, pooled over many learners. With one user there is no interaction matrix, so the question becomes what this learner would want next, from three signals kept apart: topical relevance, difficulty and novelty. Design notes only so far.

Next A first prototype (embed senses, cluster, propose, log every decision), evaluated within one subject as a crossover design.

  • recommender systems
  • cold start
  • information retrieval
  • language learning
  • llm
  • nlp
  • sentence embeddings
  • topic modeling

Agents and context engineering

3 projects

Preparing content so that an LLM agent can use it: extraction to clean Markdown, retrieval done ahead of the agent, and the token cost measured.

Agentic Org Planner

active

A planning system in plain-text Org files in which an LLM agent is a working participant. The agent runs beside the agenda, knows the file and selection in front of you and, from a generated context file, the conventions of your plan; it adds and edits tasks and files what you captured during the day, showing each change as a diff to accept or reject. Underneath is a complete Emacs configuration: a visual agenda, schedule view, capture and saved searches.

Next The audit and route skills, which review the plan and file inbox items, moving into the repo.

  • emacs
  • org mode
  • ai agents
  • claude code
  • context engineering
  • planning
  • productivity

Agent context pipeline

active

A self-hosted pipeline between what I come across during the day and the agent that files it. It captures from a phone, a browser or the command line, extracts and cleans the content, resolves links and titles, and looks up related notes ahead of the agent (BM25 recall, a cross-encoder rerank, an abstain gate). The agent's token cost with and without that lookup is measured in the design notes.

Next Cleaned of personal data and published.

  • context engineering
  • ai agents
  • information extraction
  • pdf extraction
  • ocr
  • speech to text
  • retrieval
  • self hosted

Agentic paper library

planned

A way to read and organise research papers with a coding agent at the centre. Papers from arXiv, PDFs and web articles are converted to Markdown with equations, figures and page-numbered headings, so the agent can read across the whole library and cite by section and page. Skills add a paper and place it in the topic tree, keep notes on each one, and write a literature review for an area.

Next Published without my own collection of papers.

  • context engineering
  • ai agents
  • research papers
  • pdf extraction
  • claude code

Applied LLM and speech systems

2 projects

Acervo

active

A self-hosted, offline-first vocabulary store. Paste a word or the sentence you met it in; the entry is written for you and waits in an inbox until you approve it. Around the entry, models generate a picture for each sense and spoken audio, select clips of native speakers using the word on YouTube, and turn a handful of words into audio loops and illustrated stories. Where a model is trusted and where it is checked was decided by experiments kept in the repo.

Next The quality track for articles, stories and clip selection.

  • llm
  • language learning
  • dictionary
  • multimodal
  • llm evaluation
  • text to speech
  • image generation
  • offline first
  • self hosted

LexiBeat

active

Makes loops: a handful of words, each spoken in the language you are learning and followed by its translation, over a calm music bed. Speech lands on a known beat grid and the music is procedural (numpy and scipy over CC0 samples), so a loop is byte-identical from a seed; the sound library was chosen through blind A/B listening.

Next A visualiser with the words synchronised to the audio.

  • language learning
  • audio
  • text to speech
  • procedural music
  • dsp
  • python

Implemented from scratch

3 projects

Foundation models experiments

Six blocks of implementations trained end to end: representation learning, neural search, transformers and vision transformers, multimodal captioning, audio models with efficient fine-tuning, and RLHF with PPO and preference optimisation. Begun in workshops at the Machine Learning Institute and extended independently.

  • pytorch
  • deep learning
  • transformers
  • vision transformer
  • information retrieval
  • multimodal
  • rlhf
  • from scratch

Deep RL notebooks

active

The hands-on units of the Hugging Face Deep RL course, reworked instead of copied: current library versions, the workarounds removed, and the common code moved into a small tested library. So far PPO on CartPole and Lunar Lander, and Q-learning on Frozen Lake.

Next The remaining hands-on units.

  • reinforcement learning
  • deep reinforcement learning
  • pytorch
  • jupyter notebook

Distributed training

planned

A study of distributed training in PyTorch, built up from a manual data-parallel loop.

Next DistributedDataParallel, sharded variants, and a benchmark report.

  • pytorch
  • distributed training
  • gpu

Tools and apps

8 projects

Harmonic

A minimal distraction-free Spotify playback controller for macOS. Control your music from the menu bar — skip tracks, toggle likes, and view track info at a glance.