IAlchemy Inc.
/ Applied Research Scientist Requirements · MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred) · 5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models · Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research · Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems · Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization) · Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments · Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability · Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs · Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts · Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly Preferred Skills · Hands-on experience running or supporting fine-tuning/post-training experiments (SFT, preference optimization, RLHF/RLAIF-style workflows) · Experience with multimodal evaluation (e.g., text-image, audio, video) · Experience with long-context benchmarking/evaluation and real-world context management challenges · Experience designing multi-turn, interactive, or agentic evaluation protocols · Published research and/or open-source benchmark contributions in LLM evaluation, post-training, alignment, or related areas · Experience in customer-facing applied research, technical consulting, or cross-functional product/research collaborations · Familiarity with safety, trustworthiness, and governance considerations in GenAI evaluation Application Question(s): Open to work in a night shift schedule Education: Master's (Preferred) Experience: Applied Research: 5 years (Preferred) LLM Evaluation: 5 years (Preferred) Multimodal Evaluation : 5 years (Preferred) Data Science: 5 years (Preferred) Work Location: Remote
IAlchemy Inc.