Trainium x Reactor Rolling Forcing
The AWS Neuron Science team and Reactor collaborated to bring real-time video generation to Trainium.

A kernel-centric path to real-time video generation on Trainium

Using the Neuron Kernel Interface, a Reactor–AWS collaboration tackled the dynamic shapes, memory access patterns, and cache management that make real-time autoregressive diffusion hard—building techniques that generalize across models.

In 2018, Jürgen Schmidhuber — who in 1990 was the first person to propose world models as a machine learning concept — published a paper with David Ha. They wrote about “a predictive world model” which could “extract useful representations of space and time” and use that to train agents to drive, among other things. The growth in the capability and real-world applications of world models in the eight years since that seminal paper is staggering.

The emergence of massive training datasets, paired with variational autoencoders, and both diffusion and autoregressive transformers enable today’s world models to convert text, images, audio, and video inputs into latent space that is used to build and maintain breathtaking, immersive world models which can predict state changes and respond to actions. These models have vital and expanding roles in robotics, transport, climate modeling, video game development, scientific simulation, and design prototyping.

World models can generate immersive, explorable environments that respond to user input in real time.

The sharp increase in the power and potential of world models has been accompanied by an increased need for hardware and infrastructure capable of supporting them. That is a challenge that Bryce Schmidtchen, co-founder and chief technology officer of Reactor, is well acquainted with.

“It's about real time, it's about low latency, and it's about doing that as efficiently as you can at scale,” Schmidtchen observed. Reactor is a platform that allows developers, designers, and researchers to deploy, use, and scale real-time interactive AI models. “Efficiency means everything from how you schedule the inference on the given chip, in our case Trainium, to how you think about maximally bin packing every forward pass of the model.”

The challenge of efficiency is made even more acute by Reactor’s global presence. “We think about GPU clusters in terms of regions — we have hundreds all over the world. The latent that turns into a pixel that comes off the GPU needs to be able to get to the network without hopping around through Kubernetes. And you need to do this in a way where it doesn't matter if it's a well-supported cluster or bare metal in a closet.

“And then,” Schmidtchen continued, “you need to tie this to world-class networking and media streaming that supports different codecs, different resolutions, and allows that to connect to APIs and SDKs across different languages that can support all different types of applications.”

Trainium x Reactor Rolling Forcing
Reactor's platform provides APIs and SDKs across multiple languages and codecs, supporting real-time interactive AI applications — from gaming and robotics to film and research — on AWS.

The global reach of AWS already helps Reactor to achieve many of its goals. “Everything they have at the inference layer, the level of scale that we're able to achieve in different regions is what makes a platform of this scale possible and reliable,” Schmidtchen observed.

Now, a recent collaboration between Reactor and the Amazon Neuron Science team has forged a path to even greater efficiency for world models. This is a look at the optimization strategy those teams pursued, and how that laid the foundation for future models to achieve real-time capability on Trainium.

The rise of autoregressive diffusion models

“The Neuron Science team explores new techniques for generative AI model enablement optimization,” explained Jun Wu, a principal applied scientist on the Neuron team. “This might be a new model architecture or new algorithm for optimization or a new way to generate and optimize the code for running those models on Trainium. We identify opportunities to build a prototype and make it feasible to be ported to the production pipeline.”

The team spotted one such opportunity around the usage of diffusion models in video generation. “We noticed an evolution from generating shorter, fixed-length videos to infinite-length or dynamic-length videos,” Wu explained. “Generating high-quality video frames at those lengths gave rise to the usage of autoregressive diffusion models.”
Those types of models, which combine the sequential next-token prediction found in large language models with the iteration and refinement of diffusion models, are essential for users who want to generate video where they can navigate the generated environment.

“Autoregressive diffusion means a user keystroke can be absorbed as input to the model and, conditioned on the previously generated video frame, the model can correctly decide the next move or the next scene,” Wu observed. “Most of the interactive video generation models we are seeing today are using this.”

Wu and his team, Mason Fu and Lingfan Yu, both senior applied scientists, saw a chance to optimize how those models are deployed for real-time video generation on Trainium, and saw the opportunity to do this in collaboration with Reactor. “We had been working on real-time video and interactive video generation for over 18 months at that point,” Schmidtchen noted. The teams considered various video generation techniques and aligned on utilizing Rolling Forcing, citing both its ability to consistently generate high-quality 30-second videos and its relative size.

The hard part with these models isn't quality, it's that they have to run in real time. Unlike traditional video generation, which renders a full clip offline and returns it later, models like Rolling Forcing are streaming. Each frame is generated and immediately consumed, shown to a user or fed back in as the next input. That is how developers and customers actually use them, and a frame that arrives late breaks the experience, because generation has to stay ahead of the playback timeline. “High-quality generation diffusion models pose the challenge of a very long sequence, which requires a lot of memory consumption,” Wu explained. “Rolling Forcing is relatively small, but its sequence length is very large.” That combination of a small model, a very long sequence, and a hard frame-rate floor is what makes real time so demanding. Rolling Forcing's ability to generate 16 frames per second, the standard for video playback, meant it also met latency requirements. This is where Trainium adds value, bringing the performance and memory to sustain that long-sequence workload at speed and deliver real-time generation above 16 fps rather than merely producing good frames eventually. So the teams set about enabling Rolling Forcing, as a proxy for autoregressive diffusion video generation on Trainium.

A developer-friendly approach

The Neuron Science and Reactor teams adopted a kernel-centric, bottom-up methodology aimed at making life easier on developers. They focused on three challenges — dynamic shapes, unusual memory access patterns, and heavy cache management — that real-time video generation poses for generic compilers. Each of those challenges recurs on every forward pass, so what might be an insignificant delay on other workloads can compound into latency failures in real-time video generation.

For example, the challenges posed by dynamic shapes are partly rooted in the shifting nature of scenes in real-time video generation: Imagine a shot showing a pitcher, alone on the mound, that pans out to show the other players and then thousands of fans in the stands, all within seconds. Real-time video generation also involves frames which are generated and evicted in a continuous sliding-window process. “This means attention lengths vary across forward passes, and there are two distinct passes per window: denoising, then cache cleanup,” Wu said. “All of that makes static compilation hard.”

In addition to sliding KV cache copies, video generation workloads contain operations — rotary position embeddings (RoPE) and attention transposes — whose memory access patterns are workload-specific, making them challenging for any general-purpose compiler to fully optimize. For example, in video generation RoPE must contend with three axes (height, width, and time) rather than the single position it accounts for in LLMs. “The 3D rotary embedding interleaves odd and even elements along the innermost dimension, producing many tiny data transfers when compiled generically,” Wu explained.

Finally, because KV caching happens at every layer — each one reads and writes a rolling KV cache — high throughput is required for on-device copies. The need to read old cache contents and write new ones at every layer for every step acts as another significant drag on memory.

Using the Neuron Kernel Interface

To help solve for this additional complexity, the Neuron Science team turned to the Neuron Kernel Interface (NKI).
“NKI lets developers write compute kernels that run directly on NeuronCore hardware, with precise control over how data moves between memory and compute engines,” Wu explained. He noted that NKI gives developers a self-service path to fix hotspots directly, replacing specific bottleneck operations with hardware-tuned implementations where profiling shows that simply compiling models is insufficient.

“In the work we did, the 3D-RoPE kernel went from five seconds to 1.8 milliseconds, cache copies from 23 milliseconds to 1.9 milliseconds per layer, and attention transposes were eliminated entirely by fusing them into the attention kernel,” Wu noted. “For real-time workloads where every millisecond matters, that direct hardware access is what makes production-grade performance achievable.”

Additionally, NKI-Dev-Suite — an agent for generating NKI kernels — produced a working 3D-RoPE kernel on its first attempt. “The combined effect: the pipeline used 11 GB of high-bandwidth memory, while the standard eager-mode path ran out of memory,” Wu noted.

Hybrid sharding strategy

As established, video diffusion models produce token sequences far longer than text models under the real-time requirement. That makes the challenge of self-attention more acute. “Self-attention here operates on 23,400 query tokens attending to 32,760 context tokens, and accounts for about 70% of compute time,” Yu observed. “No single core handles this efficiently without distributing the work.”

To address this, the team used a hybrid sharding strategy entailing sequence parallelism (SP), or partitioning data sequentially, and tensor parallelism (TP), which shards tensors along a specific dimension to distribute computation across multiple devices. For certain non-self-attention parts of the module, the team utilized sequence parallelism. However, for the self-attention portions, only the hybrid approach sufficed.

“LLM attention is causal and unidirectional — each token attends only to previous tokens, the KV cache grows monotonically, and there's one attention type per layer with a single cache policy,” Wu said. Rolling Forcing, however, has two attention types per block: self-attention for spatiotemporal consistency across frames and cross-attention for text conditioning and bidirectional attention within the active window, since all frames are jointly refined from noise.
That, combined with a dual-policy KV cache (a sliding window for recent context plus a permanent attention sink for global context), two forward passes per window (denoising writes noisy KV entries, then a cache update pass overwrites them with clean values, ensuring future windows always attend to clean context), and 3D video token structure constrains how the sequence can be split across cores.

Yu explained that TP alone presents a math problem. “The WAN diffusion transformer model has 12 attention heads, but we have eight Neuron cores per chip, so it's not divisible. Using TP alone means you would have to pad, but padding wastes computation.”

SP alone, on the other hand, can break the 3D structures because, as Yu noted, “You cannot guarantee the sequence partition will be right at the boundary of a frame. Your data must be at least within the granularity of a frame, but if you partition on the frame boundary, those operations won't work.” The hybrid approach splits heads across 4 cores and sequences across 2, keeping both the math and the data layout correct. “The VAE decoder used spatial W-axis sharding, achieving a super-linear 8.25 times speedup,” Yu said.

Model structure changes

The Reactor and Neuron Science teams also optimized parts of the Rolling Forcing model code to run more efficiently on Trainium. “Basically, the model has two phases: one is diffusion, the other is the cache update,” Yu said. “Those are executed in two separate runs, but the issue is the cache update phase has significantly less computation, so if you execute it in a separate round, it has far less hardware utilization. This is because we also have to shard it, and so the high-level principle is that the less data you feed to the chip, the worse hardware utilization you have.”

The team optimized the model code so that the components each of those phases have in common were batched together.

“When we encounter components that are slightly different, we split again and then handle the different components separately. But for most of the pipeline, they are batched together,” Yu explained. “This is specifically useful for Trainium, because each instance has 16 chips and each chip has eight cores. And if you partition your computation across too many cores, each call will just have a small amount of partition computation, and that's underutilizing capacity.”

The result

After employing these, and other optimizations, Reactor and the Neuron Science teams were able to successfully generate a correct video on the first end-to-end run, utilizing Trainium to deliver real-time models. And, the teams emphasized, those results are generalizable.

Trainium x Reactor Rolling Forcing
Running Rolling Forcing on a single Trainium2 chip, the collaboration achieved real-time video generation — a result the teams say generalizes across autoregressive diffusion models.

“Rolling Forcing was the pipeline, it's a very small model,” said Yahav Biran, a principal solutions architect. “You can iterate quickly on it, but it's still a robust system end to end. It has all the complexity that you have in a robust system: the encoder, the DIT, the VAE, the decoder. So basically, if you take a more robust system, it is operating on the same building blocks.”

“We're building common techniques for models that employ autoregressive diffusion which also have a requirement for real-time interaction,“ Yu added. “We're not optimizing a single model only. We're building common techniques for supporting all models with the requirements of real-time streaming.”

The future

Schmidtchen said he is excited about the future this kind of work may enable. “In the not-so-far future, every pixel will be generated in real time, interactively,” he said. “Whole stories can be created by world models in real time: stories that react, that you can engage with, that can even change their entire landscape on the fly.”

He also noted that the work Amazon is doing, and has already done, will do a great deal to make those visions a reality.

“Trainium, is clearly showing a tremendous commitment from AWS and Amazon overall,” Schmidtchen noted. “There's a clearer roadmap of higher performance, better cost performance, more scale globally. In this future where you have real-time interactive AI that needs to be distributed at scale to consumers, physical AI, and more, Trainium is very well positioned to work very well at the inference layer—in terms of its parallelization and its memory and its software stack—and integrate nicely with AWS's global scale infrastructure. We are excited to continue working closely with Amazon as we explore the untapped potential of world models. This is just the beginning.”

Research areas

Related content

IN, KA, Bengaluru
Alexa+ is the world’s best Generative AI powered personal assistant / agent for consumers, and is becoming the conversational AI interface for Amazon services with the launch of Alexa for Shopping on Amazon.com and Amazon mobile app. At Alexa Ads, we are creating industry's first and most advanced Agentic Advertising products to drive Agentic Commerce. We are seeking an Applied Scientist to join our newly expanding team in India focused on Alexa Agentic/Conversational Ads and Personalization. In this role, you will build machine learning models that seamlessly and naturally integrate relevant advertising into the Alexa experience while deeply personalizing user interactions. You will work closely with other scientists, engineers, and product managers to take models from conception to production. Key job responsibilities - Design, develop, and evaluate innovative machine learning and deep learning models for natural language processing (NLP), recommendation systems, and personalization. - Conduct hands-on data analysis and build scalable ML pipelines. - Design and run A/B experiments to measure the impact of new models on customer experience and ad performance. - Collaborate with software development engineers to deploy models into high-scale, real-time production environments. About the team We are building a new science team in Bangalore to solve some of the most impactful problems in computational advertising. This isn't about tweaking existing models as we are rethinking how ads are ranked, priced, and personalized across voice-first and screen-first surfaces. These are problems that don't have textbook solutions. Key points to note about the team: 🧪 Greenfield team - you are not joining a mature org with rigid processes. You will shape the science roadmap, pick the problems, and define the culture from day one. 📈 Direct business impact — your models directly drive revenue. No yearly cycles to see if your work matters. 🌏 Global scope, local autonomy — collaborate with scientists and engineers across Seattle, Sunnyvale, and Bangalore, but own your problem space end-to-end. 🎓 Ship AND Publish: We encourage top-tier publications (NeurIPS, ACL, EMNLP, KDD, ICML, WWW) while ensuring your research hits production.
US, CA, Palo Alto
Are you passionate about solving big problems from ground-up? Do you enjoy building new state-of-the-art products at internet scale? Come lead the innovation in this startup team, vertical ad products. This is a green field problem without a known answer or a pattern to follow. We have ambitious vision to simplify full funnel advertising solutions, at scale, with specialized agentic AI-powered models and diversify the demand to strategic verticals including finserv, autos, locals.. etc. We are seeking an experienced Sr Data Scientist to drive innovation in our Ads Foundational Model. In this individual contributor role, you will apply advanced machine learning techniques to improve advertiser performance and customer experience. Key job responsibilities As a Data Scientist on this team, you will: 1. Develop and drive the science strategy for Ads Foundational Model (Ads-FM), aligning it with the program's objectives and overall business goals. 2. Identify high-impact opportunities within Ads-FM program and lead the ideation, planning, and execution of science initiatives to address them. 3. Build and deploy machine learning models using computer vision, natural language processing, and deep learning to evaluate and enhance ad effectiveness. 4. Develop algorithms that extract meaningful signals from image, video, and audio content to predict and improve customer engagement 5. Leverage Amazon's extensive data repository to create predictive models that generate actionable recommendations for more compelling ad creative 6. Collaborate with business leaders and cross-functional teams to implement ML-powered solutions 7. Contribute to the ML roadmap for the Ads-FM program through innovation and research.
IN, TS, Hyderabad
Are you passionate about solving complex problems with machine learning and scientific rigor? As an Applied Scientist I at Amazon, you will translate real-world business challenges into well-defined scientific problems and build solutions that directly benefit customers. You will work alongside experienced scientists and engineers, applying your expertise in areas such as natural language processing, computer vision, or robotics to design experiments, develop models, and deliver production-ready code. This is a role where your curiosity and technical depth will drive meaningful impact from day one. Key job responsibilities - Design, develop, and implement machine learning models and algorithms to solve well-defined business problems, mapping business goals and metrics to scientific approaches and evaluation criteria. - Write secure, stable, testable, and maintainable production code, applying state-of-the-art data structures and algorithms while following software development best practices at a high quality bar. - Conduct rigorous experiments to evaluate model performance, benchmark results against current research, and iterate on solutions to improve accuracy and customer outcomes. - Collaborate with team members to scope technical approaches, communicate findings through internal research reports, and contribute to peer-reviewed publications when aligned with business needs. - Stay current with research trends in your area of expertise, champion the adoption of recent scientific advancements, and help onboard and mentor scientist interns. A day in the life You might start your morning reviewing experiment results from a model you trained, analyzing performance metrics and identifying areas for improvement. After a design discussion with your team, you refine your approach and push updated code for review. In the afternoon, you read a recent research paper recommended by a senior scientist, exploring whether a new technique could improve your current solution. You wrap up by documenting your methodology so teammates can understand and build on your work. About the team Our team is focused on applying scientific methods and machine learning to solve problems that matter to Amazon's customers. We value rigorous experimentation, clear communication, and a collaborative environment where scientists at every stage of their career can grow. We are building toward solutions that push the boundaries of what is possible, and we are looking for curious, thoughtful scientists who want to contribute to that mission and learn alongside a supportive group of peers.
IN, TS, Hyderabad
Welcome to the Worldwide Returns & ReCommerce team (WWR&R) at Amazon.com. WWR&R is an agile, innovative organization dedicated to ‘making zero happen’ to benefit our customers, our company, and the environment. Our goal is to achieve the three zeroes: zero cost of returns, zero waste, and zero defects. We do this by developing products and driving truly innovative operational excellence to help customers keep what they buy, recover returned and damaged product value, keep thousands of tons of waste from landfills, and create the best customer returns experience in the world. We have an eye to the future – we create long-term value at Amazon by focusing not just on the bottom line, but on the planet. We are building the most sustainable re-use channel we can by driving multiple aspects of the Circular Economy for Amazon – Returns & ReCommerce. Amazon WWR&R is comprised of business, product, operational, program, software engineering and data teams that manage the life of a returned or damaged product from a customer to the warehouse and on to its next best use. Our work is broad and deep: we train machine learning models to automate routing and find signals to optimize re-use; we invent new channels to give products a second life; we develop highly respected product support to help customers love what they buy; we pilot smarter product evaluations; we work from the customer backward to find ways to make the return experience remarkably delightful and easy; and we do it all while scrutinizing our business with laser focus. You will help create everything from customer-facing and vendor-facing websites to the internal software and tools behind the reverse-logistics process. You can develop scalable, high-availability solutions to solve complex and broad business problems. We are a group that has fun at work while driving incredible customer, business, and environmental impact. We are backed by a strong leadership group dedicated to operational excellence that empowers a reasonable work-life balance. As an established, experienced team, we offer the scope and support needed for substantial career growth. Amazon is earth’s most customer-centric company and through WWR&R, the earth is our customer too. Come join us and innovate with the Amazon Worldwide Returns & ReCommerce team! Key job responsibilities * Design, develop, and evaluate highly innovative models for Natural Language Programming (NLP), Large Language Model (LLM), or Large Computer Vision Models. * Use SQL to query and analyze the data. * Use Python, Jupyter notebook, and Pytorch to train/test/deploy ML models. * Use machine learning and analytical techniques to create scalable solutions for business problems. * Research and implement novel machine learning and statistical approaches. * Mentor interns. * Work closely with data & software engineering teams to build model implementations and integrate successful models and algorithms in production systems at very large scale. About the team When a customer returns a package to Amazon, the request and package will be passed through our WWRR machine learning (ML) systems so that we could improve the customer experience, identify return root cause, optimize re-use, and evaluate the returned package. Our problems touch multiple modalities spanning from: textual, categorical, image, to speech data. We operate at large scale and rely on state-of-the-art modeling techniques to power our ML models: XGBoost, BERT, Vision Transformers, Large Language Models.
US, TX, Austin
Amazon Leo is an initiative to launch a constellation of Low Earth Orbit satellites providing low-latency, high-speed broadband connectivity to unserved and underserved communities around the world. As a Communication Systems Research Scientist, this role owns the research and system design of the radio resource management (RRM) and radio access layers of Amazon Leo’s direct-to-device (D2D) system, delivering 3GPP-compliant service to unmodified commercial handsets. The Role: Be part of the team defining the communication system and architecture of Amazon’s direct-to-device wireless network and analyzing its system level performance: beam and cell capacity, spectral efficiency, coverage, latency and service availability. This is a unique opportunity to innovate with few legacy constraints, in a segment where the standard itself is still being written. This role leads the research and system design of radio resource management (RRM) for a 3GPP Non-Terrestrial Network (NTN), where D2D upends terrestrial assumptions: a power-limited handset with a near-isotropic antenna, very large cells, hopping beams, large time-varying delay and Doppler, and scarce shared spectrum. RRM in time, frequency and spatial domains is the focus, but the role reasons across the stack, from L1/L2 up through RRC, NAS and 5GC interworking. Agentic AI is expected to be a standard part of the work for development, optimization, tests and debugging, with the scientist accountable for the algorithms, models and conclusions. Export Control Requirement: Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum. Key job responsibilities • Research, design and specify RRM algorithms for Amazon Leo’s 3GPP-based D2D system: MAC scheduling, link adaptation, power control, HARQ strategy, DRX, admission and congestion control, and load balancing, mapping 5QI and QoS flow requirements to scheduler behavior across voice, messaging, emergency and data services. • Treat beam management as part of joint resource optimization, not a standalone process, optimizing it with band assignment, packet scheduling and user pairing in multi-user MIMO (MU-MIMO). • Define the RRM framework for NTN conditions: earth-fixed and earth-moving cells, large time-varying propagation delay, ephemeris-assisted timing and Doppler pre-compensation, extended timing advance, selective HARQ feedback disabling, feeder link and satellite handovers, and interference and spectrum sharing across beams, satellites and terrestrial networks using the same MNO spectrum. • Design mobility and service continuity for a network where the base stations (i.e., satellites) move rather than the user: idle and connected mode mobility, location and time based conditional handover, cell reselection, paging, tracking area design, and NTN-to-terrestrial continuity. • Specify supporting L1/L2 elements with the PHY team: numerology under Doppler, PRACH and initial access, coverage enhancement through repetition, synchronization at low SNR, receiver abstraction, and FEC and BLER modeling for link adaptation. • Keep the radio design coherent with the networking layers: RRC and NAS, RLC and PDCP over long-RTT links, CU/DU split, NTN gateway and 5GC/EPC integration, and transport behavior. • Develop link-level and system-level simulators capturing constellation dynamics, beam patterns, handset characteristics, traffic models and RRM behavior, and use agentic AI across that loop: build and refactor simulation code, scale parameter sweeps, optimize scheduler and link adaptation parameters, explore configuration spaces too large to sweep by hand, maintain regression tests, and triage failures across logs, traces and over-the-air captures. • Translate research into system requirements and implementation-level specifications, and work with modem, payload, ground, RF, ASIC and Testbed teams through integration, field trials and link bring-up, root-causing gaps between simulation, implementation and over-the-air behavior in a fast-paced environment. • Represent Amazon Leo in 3GPP and other standards development organizations, develop and defend contributions on NTN and D2D work items, and contribute patents and publications.
US, CA, San Diego
Amazon Leo is an initiative to launch a constellation of Low Earth Orbit satellites that will provide low-latency, high-speed broadband connectivity to unserved and underserved communities around the world. Come work at Amazon! The Role: Be part of the team defining the overall communication system and architecture of Leo’s broadband wireless network. This is a unique opportunity to innovate and define groundbreaking wireless technology with few legacy constraints. The team develops and designs the communication system of Leo and analyzes its overall system level performance such as for overall throughput, latency, system availability, packet loss etc. This role in particular will be responsible for leading the effort in integration, verification and testing of the systems especially focused on MAC and higher layer testing. This role will also be responsible developing and testing advanced L1/L2/L3 concept to improve the performance and reliability of the LEO network. This role will also be part of a team and develop simulation tools with particular emphasis on modeling the physical layer aspects such as advanced receiver modeling and abstraction, interference cancellation techniques, FEC abstraction models etc. In this role you will: - Work within a project team and take the responsibility for the Leo’s communication system design, system integration and verification. - Work as a part of the team in building a suite of system and network simulation services in Matlab / C++ / Python - Develop requirements from system level to HW/SW level and define test cases associated with the requirements. - Identify additional HW and SW that are needed for the purposes of verification and guide the HW/SW development team in the development of these test solutions/tools// - Work closely with implementation teams to simulate expected system level performance and provide quick feedback on potential improvements - Write scripts / code for functions / features required for specific simulation, testing and verification of given RF system EXPORT CONTROL REQUIREMENTS Due to applicable export control laws and regulations, candidates must be a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum.
US, WA, Seattle
Amazon Economics is seeking Structural IO Economist (STRUC) Interns who are passionate about applying structural econometric methods to solve real-world business challenges. STRUC economists specialize in the econometric analysis of models that involve the estimation of fundamental preferences and strategic effects. In this full-time internship (40 hours per week, with hourly compensation), you'll work with large-scale datasets to model strategic decision-making and inform business optimization, gaining hands-on experience that's directly applicable to dissertation writing and future career placement. By applying to this role, you are automatically being considered for all our available STRUC internships in 2027. Key job responsibilities As a STRUC Economist Intern, you'll specialize in structural econometric analysis to estimate fundamental preferences and strategic effects in complex business environments. Your responsibilities include: - Analyze large-scale datasets using structural econometric techniques to solve complex business challenges - Applying discrete choice models and methods, including logistic regression family models (such as BLP, nested logit) and models with alternative distributional assumptions - Utilizing advanced structural methods including dynamic models of customer or firm decisions over time, applied game theory (entry and exit of firms), auction models, and labor market models - Building datasets and performing data analysis at scale - Collaborating with economists, scientists, and business leaders to develop data-driven insights and strategic recommendations - Tackling diverse challenges including pricing analysis, competition modeling, strategic behavior estimation, contract design, and marketing strategy optimization - Helping business partners formalize and estimate business objectives to drive optimal decision-making and customer value - Build and refine comprehensive datasets for in-depth structural economic analysis - Present complex analytical findings to business leaders and stakeholders
US, VA, Arlington
Want to help Amazon tell its customer-centric story around the world and work in a highly cross-functional environment with economists, lawyers, scientists, public policy, public relations, and business teams? If yes, keep reading! You'll join a team of economists, engineers, and lawyers to develop economic analysis and evidence supporting legal and regulatory matters across all our lines of business worldwide—including retail, marketplace services, AWS, consumer experience, shopping and search, and operations. In this role, you will have exposure to complex regulatory issues that are of high strategic importance to the company and will develop significant expertise on the economics of Amazon’s business operations and the industries in which it operates. If you're an economist with a passion for the current legal and policy debate, strong practical judgment and creative problem-solving skills, a love of communicating economic ideas to non-technical audiences, a knack for distilling data and economic models into key insights, and a track record of delivering results fast, we want to talk to you! Key job responsibilities • Provide data-driven guidance on high-stakes legal and regulatory questions facing Amazon worldwide • Collaborate with economists, scientists, engineers, and non-technical partners on high-impact projects with global scope • Partner with global public policy teams to apply economic analyses to current policy debates on competition, AI, and related issues • Engage with external stakeholders to drive deeper understanding of Amazon’s business model and the value it develops for the economy • Support requests for economic analyses and data in ongoing regulatory and litigation matters worldwide • Synthesize business facts and data into compelling economic narratives, translating complex findings into actionable insights • Advise stakeholders across Amazon on a broad spectrum of complex and often novel economic issues • Conduct, direct, and coordinate all phases of research projects—defining key questions, evaluating methodology, executing analysis, and communicating results
US, WA, Seattle
Amazon Economics is seeking Reduced Form Causal Analysis (RFCA) Economist Interns who are passionate about applying econometric methods to solve real-world business challenges. RFCA represents the largest group of economists at Amazon, and these core econometric methods are fundamental to economic analysis across the company. In this a full-time internship (40 hours per week, with hourly compensation). You'll work with large-scale datasets to analyze causal relationships and inform strategic business decisions, gaining hands-on experience that's directly applicable to dissertation writing and future career placement. By applying to this role, you are automatically being considered for all our available RFCA internships in 2027. Key job responsibilities As an RFCA Economist Intern, you'll specialize in econometric analysis to determine causal relationships in complex business environments. Your responsibilities include: - Analyze large-scale datasets using advanced econometric techniques to solve complex business challenges - Applying econometric techniques such as regression analysis, binary variable models, cross-section and panel data analysis, instrumental variables, and treatment effects estimation - Utilizing advanced methods including differences-in-differences, propensity score matching, synthetic controls, and experimental design - Building datasets and performing data analysis at scale - Collaborating with economists, scientists, and business leaders to develop data-driven insights and strategic recommendations - Tackling diverse challenges including program evaluation, elasticity estimation, customer behavior analysis, and predictive modeling that accounts for seasonality and time trends - Build and refine comprehensive datasets for in-depth economic analysis - Present complex analytical findings to business leaders and stakeholders
US, WA, Bellevue
FBA AI Science and Analytics accelerates the AI-native transformation of Fulfillment by Amazon by building, integrating, and scaling AI-powered data & science products and seller-facing experiences that drive operational efficiency and growth across Fulfillment by Amazon globally. We learn seller behaviors, design the policies and incentives that shape their experience, and ship science products that help third-party sellers grow topline and cut operating costs at Amazon scale. Our work sits at the intersection of machine learning, statistics, economics, operations research, and GenAI/LLMs. We're looking for a Senior Applied Scientist who wants to put GenAI to work on a hard, high-visibility problem: building next-generation multi-agent systems that interact with millions of sellers and guide them through their toughest challenges at scale. You'll own solutions spanning supervised and unsupervised learning, recommendation systems, statistical learning, LLMs, harness engineering, and reinforcement learning. The ambition is to make AI a native layer in every seller decision rather than a separate tool sellers must adopt, delivering actionable insight in minutes, not days. You'll shape end-to-end experiences across the highest-frequency seller workflows, including inventory optimization, inbound efficiency, defect improvements, reimbursements, and capacity planning. The role carries direct visibility with senior Amazon business leaders and works together with fellow scientists, engineers, and product teams to launch production-grade agentic capabilities. Key job responsibilities - Design, build and deploy FBA’s GenAI architectures end to end. - Apply state-of-the-art ML and GenAI solve diverse business problems across seller supply chain systems. - Define the team’s long-term science vision and roadmap, driven fundamentally from our customers' needs, translating those directions into specific plans for scientists, engineers, and product partners. - Partner closely with scientists and software engineers to drive real-time model implementations and deliver high-impact features. - Establish scalable, efficient, automated processes for large scale data analyses, model benchmarking, model evaluation and model implementation. - Advocate the right ML solutions to business stakeholders, engineering teams, as well as executive level decision makers