About The Opportunity We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt
About The Opportunity We are looking for experienced Government professionals to support a cutting-edge AI benchmarking project focused on Indonesian public sector environments. As a Subject Matter Expert, you will create and review realistic, high-quality scenarios
About The Opportunity We are looking for experienced Retail professionals to support a cutting-edge AI benchmarking project focused on Indonesian retail and consumer-facing environments. As a Subject Matter Expert, you will create and review realistic, high-quality