
LLM Evaluation Intern
About this role
This listing aggregates 18 active LLM (Large Language Model) evaluation internship opportunities posted on Internshala, India's leading platform for student and early-career internship placements. These roles are designed for candidates interested in the fast-growing generative AI space, with core responsibilities centered on assessing LLM performance, accuracy, safety, and alignment across use cases including conversational AI, content generation, code assistance, and multilingual model testing.
Interns will collaborate with AI researchers, data scientists, and product teams to design and implement custom evaluation frameworks, run standardized and ad-hoc benchmark tests, identify model hallucinations, biases, or factual errors, and document actionable insights to improve model performance. Most internships are 3 to 6 months in duration, with flexible work options including remote, hybrid, and in-office arrangements depending on the hiring employer. Monthly stipends range from ₹10,000 to ₹35,000, scaled based on the employer, intern qualifications, and work location.
These roles are ideal for students and recent graduates pursuing degrees in computer science, AI/ML, computational linguistics, data science, or related technical fields, looking to build practical, portfolio-worthy experience in LLM evaluation and launch a career in generative AI.
Requirements
- Currently enrolled in or recently graduated from a bachelor's or master's program in Computer Science, Artificial Intelligence/Machine Learning, Computational Linguistics, Data Science, or a related technical discipline
- Foundational understanding of large language models, generative AI concepts, and common LLM evaluation benchmarks (e.g., MMLU, HellaSwag, GSM8K) is preferred
- Strong analytical skills and attention to detail to identify model errors, hallucinations, biases, or factual inconsistencies
- Basic proficiency in Python and familiarity with data analysis or annotation tools is a plus
- Excellent written and verbal communication skills to document evaluation findings and collaborate with cross-functional teams
- Prior experience with AI model testing, prompt engineering, or NLP projects is advantageous but not required for most listed roles
Quick Info
Employment Type
Internship
Location
Multiple locations across India (remote, hybrid, and in-office options available)
Department
Artificial Intelligence / LLM Evaluation
You'll leave Agentic Academy to apply on the employer's site. We curate these listings for visibility only.