AnmeldenApp laden

Jobs › Stellenangebot vor 3 Tagen

ARTIFICIAL ANALYSIS | MTS, Forward Deployed Engineer + more | San Francisco, CA (ON-SITE, 5 days) or Melbourne, Australia (ON-SITE, 3 days) | Full-time | https://artificialanalysis.ai/careers

ARTIFICIAL ANALYSIS · San Francisco, CA (ON-SITE, 5 days) or Melbourne, Australia (ON-SITE, 3 days)

Vor OrtEnglisch

Über die Stelle

Aus der Anzeige des Arbeitgebers · ARTIFICIAL ANALYSIS · veröffentlicht am 1. Oktober 2026

ARTIFICIAL ANALYSIS | MTS, Forward Deployed Engineer + more | San Francisco, CA (ON-SITE, 5 days) or Melbourne, Australia (ON-SITE, 3 days) | Full-time | https://artificialanalysis.ai/careers We're an independent AI benchmarking company. We benchmark language models, inference providers, hardware, image/video and speech, and our Intelligence Index is widely cited when new models launch. Team of ~50, backed by Nat Friedman, Daniel Gross, Andrew Ng and others. The people who build our evals also work directly with the AI labs, often on pre-release models. Flat structure, lots of ownership. Main hiring needs:

  • Forward Deployed Engineer (FDE), Language Models: run our LLM benchmarking stack and work directly with the labs
  • Member of Technical Staff (MTS), Language Model Evaluations: build frontier evals (datasets, harnesses, benchmarks)
  • MTS, Inference: benchmark speed, quality and price across serverless inference providers
  • MTS, Hardware: benchmark GPUs, TPUs and custom silicon
  • MTS, Speech: own our TTS, STT and speech-to-speech evals

Also hiring for media generation, full stack, ML engineering, robotics and product roles. Apply at https://artificialanalysis.ai/careers .

Sieh deinen Match-Score für jede Stelle

BabZituna bewertet jede Stelle nach sechs echten Kriterien gegen dein Profil und zeigt dir, WARUM sie so abschneidet, auf Fairness geprüft (das öffentliche Bias-Audit lesen).

App laden → ✓ 100 % kostenlos für Jobsuchende
Wie eine Stelle abschnitt Beispiel
Fähigkeiten96Erfahrung90Standort84Arbeitsweise74Anstellungsart61GehaltKeine Daten

Beispielwerte, kein echter Kandidat. Jedes Kriterium wird aus deinem eigenen Profil mit bis zu 100 Punkten bewertet, und eines, das wir nicht messen können, sagt das, statt zu raten.