LLM Evaluation Internship
About this role
Join our team as an LLM Evaluation Intern and dive into the cutting‑edge world of large language model assessment. In this role, you will design and implement comprehensive evaluation pipelines to measure model performance on tasks such as text generation, summarization, question answering, and reasoning. You will work with a variety of benchmark datasets, configure hyper‑parameters, and run large‑scale experiments to compare different model architectures. Your responsibilities include analyzing output quality, identifying failure modes, and producing detailed reports that guide model improvement. You will collaborate closely with research scientists and software engineers, contributing to the development of robust metrics and automated testing frameworks. The internship is ideal for motivated students or recent graduates with a strong foundation in machine learning, statistics, or computer science, and proficiency in Python and common ML libraries. We offer a flexible, remote work arrangement, a duration of six months, and a monthly stipend ranging from ₹7,500 to ₹10,000. This position provides valuable hands‑on experience in AI research and a chance to shape the future of language model evaluation.
Requirements
- Strong understanding of machine learning fundamentals
- Proficiency in Python and libraries such as TensorFlow or PyTorch
- Experience with data analysis and visualization tools
- Ability to design and run experiments
- Excellent communication and teamwork skills
Quick Info
Employment Type
Internship
Location
Remote
Department
Artificial Intelligence
You'll leave Agentic Academy to apply on the employer's site. We curate these listings for visibility only.