Roberto Segura Sala
Hi, I'm Roberto.
I'm from Cuenca and I have two small kids, so most of my time away from the keyboard goes to them. When it doesn't, it goes to the gym or to the countryside.
I'm curious by nature. Most of what I know I picked up because I wanted to understand how something worked, and I still learn best by building it myself — that's what the personal projects on this site are.
At Audiense I'm on the data platform team. I look after the architecture for accessing our data, the real-time pipelines that keep it current, and the ML processes we use to enrich it. The details are on the projects page.
Selected work
Three problems I've worked on at Audiense
Keeping a 15 TB lake current in real time
Migrated social-relationship data to Iceberg while 20+ consumers kept reading it, and finished the move from batch to real-time enrichment.
Iceberg · Kafka · Trino Applied MLInterest classification: trusting agreement instead of a score
444 categories, confident mistakes. A LoRA-tuned bi-encoder proposes, a cross-encoder reranks, and only categories both rank in their top 5 survive. Human-judged strict precision 36% → 52%.
LoRA · Cross-encoders · vLLM PlatformOne way in to the data instead of a copy per team
VEGA is an asynchronous query API over the lake. Product teams dropped their local RDS and Mongo copies, and the ETL and Spark jobs that fed them.
FastAPI · Redis · RedpandaOutside work
Things I build to understand them
- Two models, one AND — a write-up of the interest classifier above, in progress for the Audiense engineering blog.
- Hermes Expense Tracker — shared household expenses over Telegram; a FastMCP server, SQLite, one assistant per person.
- Voice agent prototypes — STT → LLM → TTS on LiveKit, with explicit state machines and a validation harness.
- AI research wiki — an Obsidian vault on agents, voice pipelines and evaluation, maintained with Hermes Agent.