Open to work - AI Evaluation and ML Engineer roles

Pathi Manikanta

Building and evaluating frontier AI, with proof.

AI Evaluation Specialist, ML Engineer and RLHF/RLAIF contributor. I audit agentic tool-use trajectories, design pass/fail rubrics and pytest verifiers, and build end-to-end ML systems. I care about finding where a model breaks - and proving it with evidence.

Pathi Manikanta portrait
Pathi Manikanta
AI Evaluation / ML Engineer / RLHF-RLAIF
97.8%
Best model accuracy shipped
10+
Projects shipped
4 yrs
Trading experience (crypto, equities)
About

Evaluation, engineering, and taxation

AI Evaluation Specialist, ML Engineer and Full Stack Developer from Andhra Pradesh, India.

B.E. Computer Science and Engineering (Data Science), Anna University (2025). At Scale AI (Outlier) I reviewed frontier AI and was promoted from Attempter to Reviewer. I designed binary pass/fail rubrics backed by pytest verifiers and audited agentic tool-use trajectories. That engagement has concluded and I am now open to new roles.

I work across AI evaluation, LLM engineering, RLHF/RLAIF and full-stack ML. My focus is finding the exact turn where a model breaks and building the deterministic verifier that proves it.

Languages: Telugu, English, Hindi, Tamil

Finance and Taxation

US Enrolled Agent (EA) credential In progress
Indian income-tax returns ITR-1 to ITR-4
4 years trading experience across crypto and Indian equities
Ex-MCX Advisor at Zebu Share and Wealth Management
Skills and Expertise

The full stack of building and proving

From rubric design and trajectory review to ML systems, web engineering, and taxation.

AI Evaluation and ML
RLHF/RLAIFRubric designAgentic trajectory review Red teamingPrompt engineeringHallucination detectionMultimodal assessment Pythonpytestscikit-learn TensorFlow/KerasSHAPNLTK
Engineering and Cloud
ReactJavaScriptSQL GitAWS
Finance and Taxation
US EA (in progress)ITR-1 to ITR-4 4 yrs tradingEx-MCX Advisor
Frontier AI Evaluation

Pipelines, rubrics, and verifiers

Golden/silver trajectory pipelines, safety gates, and QC tooling for frontier models.

OpenClaw Atlas

Private
Scale AI (Outlier)

Golden/silver memory-trap pipeline for authoring and QC-ing agentic RL trajectories: story -> discrepancy -> failure grade -> silver fix -> pytest verifiers -> QC.

Blue Shell

Private
Scale AI (Outlier)

Two-turn golden/silver capabilities-RL pipeline across six tools where Turn 1 must fail a rubric trap and a Turn-2 nudge must pass 100 percent.

Lobster Safety

Private
Scale AI (Outlier)

Nine-worker safety-RL pipeline with a hard PASS/STOP gate, grading trajectories against an SSOT taxonomy (F1-F10 failure types, S0-S3 severity).

TrajLens

Public
Personal, 2026

Interactive agentic-trajectory auditor: grades 14 recorded tool-use traces against a 9-criterion rubric of deterministic verifiers (incl. safety checks for buried errors and unauthorized actions), and measures an LLM-judge against human labels with Cohen's kappa. Paste your own trajectory JSON and grade it live.

JudgeLab

Public
Personal, 2026

LLM-as-a-judge reliability lab: measures GPT-4-as-a-judge against 3,355 real human expert votes from the MT-Bench study. Cohen's kappa 0.46 on decisive verdicts, a 15.8% order-swap inconsistency rate (position bias), a quantified verbosity preference, and per-category reliability showing the judge collapses on writing tasks. Real public data, no API key.

EvalGate

Public
Personal, 2026

CI/CD pipeline that blocks bad models (GitHub Actions): every push runs 11 tests, retrains the model, enforces a hard quality gate (ROC-AUC threshold plus a CV-vs-test drift check), promotes the gated artifact between workflows, and auto-deploys a live model-quality dashboard to GitHub Pages. Weekly scheduled retrains, deployment environments, zero manual steps.

Project Seal - Authoring

Private
Handshake AI

Runbook for writing hard single-answer multi-hop search questions that defeat frontier models; 8 quality rules, 18 techniques.

Project Seal - QC Checker

Public
Handshake AI

Automated QC checker that runs a task package against the rejection rules.

ML and Engineering Projects

Shipped, end to end

HandPilot

Computer Vision, 2026

Webcam gesture control for Windows: MediaPipe and OpenCV hand tracking drive the pointer, pinch clicks, scrolling, and an on-screen keyboard, with optional object and pen following. PySide6 overlay, Windows CI tests.

PythonMediaPipeOpenCVPySide6

Chatbox

Full-stack, 2026

Streaming chat app with a pluggable model backend: FastAPI and Server-Sent Events on the server, React and TypeScript in the browser. Ships an offline replay provider so it runs with no API key; also supports Gemini and Claude. CI included.

FastAPISSEReactTypeScript

Customer Churn Predictor

ML & Data, 2026

End-to-end ML app predicting telecom churn (ROC-AUC 0.843) on 7,000+ customers; Gradient Boosting; Streamlit; live risk scoring.

Pythonscikit-learnStreamlit

Diabetes Foot Pressure Detection

ML & Data, 2025

Full ML pipeline from Arduino sensors to classifiers at 92.5% accuracy; caught and fixed session-level data leakage with GroupKFold.

ArduinoMATLABscikit-learnKeras

Spam Email Classifier

NLP, 2024

NLTK + TF-IDF, 97.8% accuracy with Linear SVM on 1000+ emails.

PythonNLTKTF-IDFSVM

Face Recognition (OpenCV)

Computer Vision, 2024

Real-time detection and recognition at 15+ FPS, 70%+ confidence, Haar Cascades + LBPH.

PythonOpenCVLBPH

IoT Environment Monitor

Hardware, 2023

Arduino DHT11/BMP180 streaming temperature, humidity, pressure every 2 seconds to a remote dashboard.

ArduinoDHT11BMP180

Google Doc Grid Parser

ML & Data

Parses coordinate triples from a published Google Doc and prints the hidden block-letter message.

PythonBeautifulSoup
Experience

Where I have worked

AI Evaluation Specialist, Reviewer - Scale AI (Outlier)

2026 - Engagement complete

Promoted from Attempter to Reviewer. Designed binary pass/fail rubrics backed by pytest verifiers and audited agentic tool-use trajectories on frontier AI.

Software and ML Engineer - Tensium

2026 - Present, Remote

Author CPU-based long-horizon ML tasks requiring multistep reasoning and code generation; develop, debug, and refactor Python pipelines for maintainability and performance.

AI Evaluation Contributor - Handshake AI

2026

Authored hard single-answer multi-hop search questions and built automated QC tooling against rejection rules.

MCX Advisor - Zebu Share and Wealth Management

2025 - 2026

4 years of trading experience across crypto and Indian equities. Supported client trading operations: order processing, position monitoring, and profit-and-loss updates.

Achievements and Awards

Selected highlights

Promoted Attempter to Reviewer in 3 months

Scale AI (Outlier)

Recognized for evaluation accuracy, guideline compliance, and consistent task quality.

Shipped 3 live AI-evaluation tools

TrajLens, JudgeLab, EvalGate

Public deployed apps for agentic-trajectory auditing, LLM-judge reliability, and CI/CD model gating.

CI/CD pipeline green for months

EvalGate

Automated weekly retraining and quality-gated deployments running unattended on GitHub Actions.

Caught a data-leakage bug others missed

Diabetes classification project

Traced inflated 99% accuracy to session leakage and corrected it to a truthful 92.5% with GroupKFold.

HackerRank certified

Problem Solving (Intermediate) and Frontend Developer (React)
Education and Certifications

Foundations and credentials

B.E. CSE (Data Science)

Jansons Institute of Technology (Anna University), 2021-2025 - CGPA 8.0

Higher Secondary (XII) MPC

Sri Sai Junior College, Kanigiri AP, 2021

Secondary (X)

Nagarjuna Model School, Kadapa AP, 2019 - CGPA 9.8
Certifications
Claude Code 101, Anthropic (2026)
AI Fluency for Builders, Anthropic (2026)
Introduction to Agents, Mercor (2026)
Prompt Engineering (OutlierEDU), Outlier AI (2026)
Problem Solving (Intermediate), HackerRank
Frontend Developer (React), HackerRank
AWS Academy Cloud Foundations (2022)
Microsoft 365 Productivity Advanced (2022)
Ethical Hacking and Penetration Testing, Scholiverse
Open to work

Let us build and prove the next model

Open to AI Evaluation, LLM Engineering, RLHF/RLAIF and ML Engineer roles. Remote or relocation. Fastest reply by email.