AI Evaluation Specialist, ML Engineer and RLHF/RLAIF contributor. I audit agentic tool-use trajectories, design pass/fail rubrics and pytest verifiers, and build end-to-end ML systems. I care about finding where a model breaks - and proving it with evidence.
AI Evaluation Specialist, ML Engineer and Full Stack Developer from Andhra Pradesh, India.
B.E. Computer Science and Engineering (Data Science), Anna University (2025). At Scale AI (Outlier) I reviewed frontier AI and was promoted from Attempter to Reviewer. I designed binary pass/fail rubrics backed by pytest verifiers and audited agentic tool-use trajectories. That engagement has concluded and I am now open to new roles.
I work across AI evaluation, LLM engineering, RLHF/RLAIF and full-stack ML. My focus is finding the exact turn where a model breaks and building the deterministic verifier that proves it.
Languages: Telugu, English, Hindi, Tamil
From rubric design and trajectory review to ML systems, web engineering, and taxation.
Golden/silver trajectory pipelines, safety gates, and QC tooling for frontier models.
Golden/silver memory-trap pipeline for authoring and QC-ing agentic RL trajectories: story -> discrepancy -> failure grade -> silver fix -> pytest verifiers -> QC.
Two-turn golden/silver capabilities-RL pipeline across six tools where Turn 1 must fail a rubric trap and a Turn-2 nudge must pass 100 percent.
Nine-worker safety-RL pipeline with a hard PASS/STOP gate, grading trajectories against an SSOT taxonomy (F1-F10 failure types, S0-S3 severity).
Interactive agentic-trajectory auditor: grades 14 recorded tool-use traces against a 9-criterion rubric of deterministic verifiers (incl. safety checks for buried errors and unauthorized actions), and measures an LLM-judge against human labels with Cohen's kappa. Paste your own trajectory JSON and grade it live.
LLM-as-a-judge reliability lab: measures GPT-4-as-a-judge against 3,355 real human expert votes from the MT-Bench study. Cohen's kappa 0.46 on decisive verdicts, a 15.8% order-swap inconsistency rate (position bias), a quantified verbosity preference, and per-category reliability showing the judge collapses on writing tasks. Real public data, no API key.
CI/CD pipeline that blocks bad models (GitHub Actions): every push runs 11 tests, retrains the model, enforces a hard quality gate (ROC-AUC threshold plus a CV-vs-test drift check), promotes the gated artifact between workflows, and auto-deploys a live model-quality dashboard to GitHub Pages. Weekly scheduled retrains, deployment environments, zero manual steps.
Runbook for writing hard single-answer multi-hop search questions that defeat frontier models; 8 quality rules, 18 techniques.
Webcam gesture control for Windows: MediaPipe and OpenCV hand tracking drive the pointer, pinch clicks, scrolling, and an on-screen keyboard, with optional object and pen following. PySide6 overlay, Windows CI tests.
Streaming chat app with a pluggable model backend: FastAPI and Server-Sent Events on the server, React and TypeScript in the browser. Ships an offline replay provider so it runs with no API key; also supports Gemini and Claude. CI included.
End-to-end ML app predicting telecom churn (ROC-AUC 0.843) on 7,000+ customers; Gradient Boosting; Streamlit; live risk scoring.
Full ML pipeline from Arduino sensors to classifiers at 92.5% accuracy; caught and fixed session-level data leakage with GroupKFold.
NLTK + TF-IDF, 97.8% accuracy with Linear SVM on 1000+ emails.
Real-time detection and recognition at 15+ FPS, 70%+ confidence, Haar Cascades + LBPH.
Arduino DHT11/BMP180 streaming temperature, humidity, pressure every 2 seconds to a remote dashboard.
Parses coordinate triples from a published Google Doc and prints the hidden block-letter message.
Promoted from Attempter to Reviewer. Designed binary pass/fail rubrics backed by pytest verifiers and audited agentic tool-use trajectories on frontier AI.
Author CPU-based long-horizon ML tasks requiring multistep reasoning and code generation; develop, debug, and refactor Python pipelines for maintainability and performance.
Authored hard single-answer multi-hop search questions and built automated QC tooling against rejection rules.
4 years of trading experience across crypto and Indian equities. Supported client trading operations: order processing, position monitoring, and profit-and-loss updates.
Recognized for evaluation accuracy, guideline compliance, and consistent task quality.
Public deployed apps for agentic-trajectory auditing, LLM-judge reliability, and CI/CD model gating.
Automated weekly retraining and quality-gated deployments running unattended on GitHub Actions.
Traced inflated 99% accuracy to session leakage and corrected it to a truthful 92.5% with GroupKFold.
Open to AI Evaluation, LLM Engineering, RLHF/RLAIF and ML Engineer roles. Remote or relocation. Fastest reply by email.