While 2023 was mostly focused on training foundation models for AI, companies are expected to begin performing inference at scale on their models in 2024.
Grader A: hit
Grader B: hit
agrees with A, without seeing its verdict
Audit: confirmed
in the fixed audit sample (D11)
Published: hit
Grader A: hit
Google processed 9.7 trillion tokens a month in May 2024 (480 trillion a year later), i.e. inference at very large scale in 2024. Nvidia already estimated ~40% of fiscal-2024 (Feb 2023-Jan 2024) data-center revenue was for inference, so 'begin' occurred partly before 2024; under D30's early-event rule this still counts because the state also held, and grew, in 2024.
Test: State of affairs (D30): large-scale production inference in 2024, shown by inference volumes or inference share of AI compute revenue.
Grader B: hit
Nvidia estimated that about 40% of its Data Center revenue in fiscal 2024 (to Jan 2024) was for AI inference, and over 40% for the trailing four quarters to Q2 FY2025 (mid-2024). Inference at scale held in 2024 (it was already under way, but D30 early occurrence counts because it also holds in 2024) -> hit.
Test: D30 state of affairs: inference is a large share of AI compute demand during 2024.