Language mastery

In two years of engineering, I've mastered every programming language.

Obviously, that is not true.

But I do build quickly, learn fast, and use AI well enough to make a slightly ridiculous claim feel almost believable. Before I show you the highly scientific evidence, here is a little about me.

So, who is behind these suspiciously good numbers?

I'm San, an engineer with a background in data who builds with AI.

I do not see AI as a shortcut around engineering, but as leverage: a way to explore ideas sooner, move through early versions faster, and focus more attention on the problems that matter.

I follow it closely because the pace of progress still feels unreal. New models and tools keep expanding what engineers can create, and I am constantly surprised by the gap between what once felt out of reach and what is now possible.

Before engineering, I worked in financial services, where I became interested in the systems behind everyday work: how information moves, where processes become difficult, and what better tooling can change. I later completed the Northcoders Data Engineering bootcamp and began building event-driven services, pipelines, and practical data tools.

Today, I work across data and AI while running Little Lotus, a web design and AI automation studio, a place to build useful things, explore new ideas, and stay close to a field that keeps moving.

Right, now onto the benchmarks.

Evaluated against the same suites as the frontier models

Humbly, I subjected myself to the same rigorous evaluations used to measure the world's leading language models. The results are in. I've decided to make them public.

san-fernandes / eval-results

AIME 2026

Competition-grade mathematics

0.0%

matches GPT-5, who also got a perfect score

GPQA Diamond

PhD-level science, under pressure

0.0%

coincidentally identical to Claude Mythos Preview

MMLU

Knowing things, in general

0.0%

exactly where the frontier models cluster

HumanEval

Code that runs on the first try

0.0%

suspiciously close to GPT-5.3 Codex

SWE-bench Verified

Fixing real bugs in real repositories

0.0%

the current Claude Opus number, give or take

Methodology: vibes, mostly. Scores are self-reported and may bear an uncanny resemblance to the early-2026 frontier-model leaderboard. Peer reviewed by zero peers. No San was fine-tuned in the making of this chart.

Need someone who ships at the speed of a frontier model?

I'm open to data and AI engineering roles, and Little Lotus is taking on projects. Say hello. I reply faster than I benchmark.