NOW BUZZING 헤어컬러가 브라운과 오렌지로 쪼개지는 이… · 코치, 85주년에 '낡음'을 미래로 바꾸… · 프레피룩, 왜 더 플라자에서 다시 태어났…◆
NOW BUZZING 헤어컬러가 브라운과 오렌지로 쪼개지는 이… · 코치, 85주년에 '낡음'을 미래로 바꾸… · 프레피룩, 왜 더 플라자에서 다시 태어났…◆
JELIBI
⌕ SUBSCRIBE
SEOUL · DAILY · ISSUE No.327 WE TRIED IT · THE PICK · BEHIND IT 2026.09.30 · WED
JELIBI
⌕ SUB
Beauty Body Gadget Bites Buzz World
HOME/GADGET/BEHIND IT
GADGET BEHIND IT · 5 MIN READ

NVIDIA, AI 추론 소프트웨어 Dynamo 공개

JD 젤리비 편집국 · 2025.07.03 SHARE · COPY LINK

무슨 발표인가

  • Triton Inference Server의 후속 오픈소스 추론 소프트웨어
  • 수천 개 GPU 규모의 추론 통신 오케스트레이션·가속화
  • 처리·생성 단계를 분리해 각 단계별 독립 최적화 가능

현재 이용 가능

원문 (영어)

GTC— NVIDIA today unveiled NVIDIA Dynamo , an open-source inference software for accelerating and scaling AI reasoning models in AI factories at the lowest cost and with the highest efficiency. Efficiently orchestrating and coordinating AI inference requests across a large fleet of GPUs is crucial to ensuring that AI factories run at the lowest possible cost to maximize token revenue generation.

As AI reasoning goes mainstream, every AI model will generate tens of thousands of tokens used to “think” with every prompt. Increasing inference performance while continually lowering the cost of inference accelerates growth and boosts revenue opportunities for service providers.

NVIDIA Dynamo, the successor to NVIDIA Triton Inference Server , is new AI inference-serving software designed to maximize token revenue generation for AI factories deploying reasoning AI models. It orchestrates and accelerates inference communication across thousands of GPUs, and uses disaggregated serving to separate the processing and generation phases of large language models (LLMs) on different GPUs.

This allows each phase to be optimized independently for its specific needs and ensures maximum GPU resource utilization. “Industries around the world are training AI models to think and learn in different ways, making them more sophisticated over time,” said Jensen Huang, founder and CEO of NVIDIA.

“To enable a future of custom reasoning AI, NVIDIA Dynamo helps serve these models at scale, driving cost savings and efficiencies across AI factories.” Using the same number of GPUs, Dynamo doubles the performance and revenue of AI factories serving Llama models on today’s NVIDIA Hopper platform.

원문: NVIDIA News — "NVIDIA Dynamo Open-Source Library Accelerates and Scales AI Reasoning Models" (2025-07-03) 공식 원문: https://nvidianews.nvidia.com/news/nvidia-dynamo-open-source-library-accelerates-and-scales-ai-reasoning-models

#NVIDIA News
THE JELIBI BRIEF

A small team that reads too much internet so you don't have to.

SUBSCRIBE →
GADGET

인스로픽, 안전장치 없이 사이버공격 자동화 모델 공개

GADGET

삼성 6개 계열사, AI 인프라 기업 헬릭스에 10억 달러 투자

GADGET

TBC, AWS와 협력해 신경세포 유래 AI 비디오 모델 출시

← GADGET BEHIND IT

NVIDIA, AI 추론 소프트웨어 Dynamo 공개

젤리비 편집국·2025.07.03·5 MIN
IN THIS PIECE
무슨 발표인가 원문 (영어)

무슨 발표인가

  • Triton Inference Server의 후속 오픈소스 추론 소프트웨어
  • 수천 개 GPU 규모의 추론 통신 오케스트레이션·가속화
  • 처리·생성 단계를 분리해 각 단계별 독립 최적화 가능

현재 이용 가능

원문 (영어)

GTC— NVIDIA today unveiled NVIDIA Dynamo , an open-source inference software for accelerating and scaling AI reasoning models in AI factories at the lowest cost and with the highest efficiency. Efficiently orchestrating and coordinating AI inference requests across a large fleet of GPUs is crucial to ensuring that AI factories run at the lowest possible cost to maximize token revenue generation.

As AI reasoning goes mainstream, every AI model will generate tens of thousands of tokens used to “think” with every prompt. Increasing inference performance while continually lowering the cost of inference accelerates growth and boosts revenue opportunities for service providers.

NVIDIA Dynamo, the successor to NVIDIA Triton Inference Server , is new AI inference-serving software designed to maximize token revenue generation for AI factories deploying reasoning AI models. It orchestrates and accelerates inference communication across thousands of GPUs, and uses disaggregated serving to separate the processing and generation phases of large language models (LLMs) on different GPUs.

This allows each phase to be optimized independently for its specific needs and ensures maximum GPU resource utilization. “Industries around the world are training AI models to think and learn in different ways, making them more sophisticated over time,” said Jensen Huang, founder and CEO of NVIDIA.

“To enable a future of custom reasoning AI, NVIDIA Dynamo helps serve these models at scale, driving cost savings and efficiencies across AI factories.” Using the same number of GPUs, Dynamo doubles the performance and revenue of AI factories serving Llama models on today’s NVIDIA Hopper platform.

원문: NVIDIA News — "NVIDIA Dynamo Open-Source Library Accelerates and Scales AI Reasoning Models" (2025-07-03) 공식 원문: https://nvidianews.nvidia.com/news/nvidia-dynamo-open-source-library-accelerates-and-scales-ai-reasoning-models

#NVIDIA News
THE JELIBI BRIEF

A small team that reads too much internet so you don't have to.

SUBSCRIBE →
MORE IN BEHIND IT
질병관리청, WHO 의료대응수단 네트워크 포럼 참석해 감염병 대응 협력 강화인스로픽, 안전장치 없이 사이버공격 자동화 모델 공개삼성 6개 계열사, AI 인프라 기업 헬릭스에 10억 달러 투자Samsung, 세계 심장의 날 캠페인으로 심장 건강 관리 강조CJ도너스캠프, 아동복지시설 교사 260명 교사교육 실시
THE JELIBI BRIEF

인터넷을 너무 많이 보는 팀이 대신 골라 옵니다.

SUBSCRIBE

바로가기 등록하시면, 더 쉽게 찾아보실 수 있습니다.