G2i
AI evaluation & benchmark engineering
Three areas of freelance work spanning task design, comparative model evaluation and the maintainability of generated code.
Read the work overviewI’m Uday, an AI engineer working across enterprise retrieval, agent workflows and Python services. These are the systems I’ve built, the problems behind them and the lessons along the way.
Professional experience across enterprise AI, automation, recommendations and applied machine learning.
Senior Developer · AI & Automation
Production RAG and agent systems
Explore 1 project & architecture notesMachine Learning Engineer
Enterprise AI and incident automation
Explore 1 project & architecture notesSoftware Engineer
Recommendations and commerce analytics
Explore 2 projects & architecture notesMachine Learning Engineer
Virtual assistants and healthcare NLP
Explore 2 projects & architecture notesTechnical Engineer · Machine Learning
Applied ML and data systems
Explore 2 projects & architecture notesSelected engagements in coding benchmarks, model evaluation, task design and computer-use data.
AI evaluation & benchmark engineering
Three areas of freelance work spanning task design, comparative model evaluation and the maintainability of generated code.
Read the work overviewSoftware-engineering task authoring
Realistic features, fixes and enhancements built around behavioral requirements, reference implementations and held-out tests in reproducible environments.
Read the work overviewLong-horizon model evaluation
Compare two models on extended build tasks, writing rubrics and rationales from their outputs and trajectories. Continue with follow-up prompts until either model exceeds 1M cumulative tokens.
Read the evaluation workflowComputer-use evaluation & interaction data
Explored computer environments and recorded interaction data for LLM training and evaluation. Reviewed interaction patterns and data quality with the team.
Read the computer-use workflowDesign an open-ended engineering prompt, then compare two models’ outputs and execution trajectories. Write rubrics and evidence-based rationales, and develop follow-up prompts that test whether each model can sustain progress over an extended session.
Read the Alignerr workflowHigh-level summaries of engineering responsibilities across these engagements.
I turn engineering concepts into visual lessons. Explore my growing collection on generative AI, agents and system design.
I’m Uday Kiran Reddy Kondreddy, an AI engineer based in Hyderabad. My work spans production retrieval systems, agent workflows, Python services, and coding-model evaluation.
I care about the part after the prototype: what breaks, how we notice, and whether our tests measure the behavior we actually need.
These project notes cover the problems I worked on, my engineering contributions and the systems behind them.
Have an interesting engineering problem? Let’s talkWorking through a retrieval problem, an agent workflow, or an evaluation that doesn’t quite add up?