We’re hiring — right now — for Research Engineers to join Polaris, a new team at Google DeepMind. As synthetic data pipelines hit their ceiling, the hardest engineering problem in AI is translating the full depth, nuance, and complexity of real-world software engineering into rigorous evaluation frameworks and training benchmarks. That’s why we’re building Polaris. Rather than generating a large number of tasks using a synthetic data pipeline, engineers on the Polaris team focus on creating a smaller number of tasks that more closely capture the complexity of real-world software engineering — taking roughly a week to complete, from ideation to final approval. Alongside building these benchmarks, you will have the autonomy to launch entirely new projects within DeepMind aimed at advancing Gemini’s training data and evaluations. You’ll create experiments, prototype implementations, design new architectures, and tackle real-world problems across AI, NLP, compilers, search, and hardware/software performance analysis all while staying connected to the wider research community through university partnerships and publishing papers. Great talent comes from anywhere. We're hiring across the entire spectrum of experience to find the best people in the world. We want to hear from you if you are: • A Competitor or Hacker: You thrive in competitive programming, math olympiads (IMO, IOI, Putnam, USAMO), or hackathons, and love constructing deeply challenging problems. • An AI-Native Builder: You already live in AI coding agents and LLM-based workflows to ship software at high velocity. • A Polyglot Engineer: You can parachute into unfamiliar programming paradigms, tools, and technical stacks and master them rapidly. If you want to build with a team at Google DeepMind, or know someone who belongs in this room, join us. Apply now → https://goo.gle/3V5Kacp
At UXV Searcher, we strongly agree that synthetic data pipelines are hitting a ceiling. The next frontier for models like Gemini requires exactly what the Polaris team is building: high-fidelity, human-driven evaluation frameworks. In our work validating complex AI search and user experience workflows, we see that capturing the nuance of real-world software engineering is the ultimate industry bottleneck right now. Brilliant move by Google DeepMind to prioritize deep, multi-day benchmark construction over shallow synthetic volume.
The ceiling is real: synthetic pipelines optimize for throughput, but the long-tail failures that matter in software engineering rarely show up in bulk-generated tasks. Curious whether Polaris is selecting its smaller task set through adversarial filtering against execution traces and repo-level constraints, or through human expert rubrics—because with MoE models, eval signal per task matters more than task count.
Why hello.
A week spent constructing one realistic task is a striking detail. The awkward dependencies and incomplete context in software projects are easy to lose in a clean benchmark. It would be fascinating to see what the team learns about evaluating those messier parts of engineering.
A small number of tasks that each take a week to build is the right trade, and it carries directly into finance. A realistic M&A or credit task can take a senior finance professional days to write, solve, and defend. Synthetic pipelines can generate more tasks, but often strip out the judgment that makes the task useful. The ceiling showing up in software is going to show up in finance just as quickly.
A week spent constructing one realistic task is a striking detail. The awkward dependencies and incomplete context in software projects are easy to lose in a clean benchmark. It would be fascinating to see what the team learns about evaluating those messier parts of engineering.
The Gemini 4 context window / reasoning improvements are genuinely useful for enterprise workflows. The governance question that doesn't get asked enough: as models get better at generating confident outputs, the gap between 'model is usually right' and 'team stopped checking' narrows. The capability curve and the accountability curve aren't moving at the same speed.
Any sort of “intelligence” probably requires less software development and more gardening. These systems are not built, they are carefully “grown” with long horizon goals in mind.