時価総額: $2.1895T 0.52%
ボリューム(24時間): $50.8664B 50.60%
  • 時価総額: $2.1895T 0.52%
  • ボリューム(24時間): $50.8664B 50.60%
  • 恐怖と貪欲の指数:
  • 時価総額: $2.1895T 0.52%
暗号
トピック
暗号化
ニュース
暗号造園
動画
トップニュース
暗号
トピック
暗号化
ニュース
暗号造園
動画
bitcoin
bitcoin

$64129.724403 USD

1.04%

ethereum
ethereum

$1892.503526 USD

-0.26%

tether
tether

$0.998974 USD

0.00%

bnb
bnb

$602.936025 USD

-0.36%

usd-coin
usd-coin

$0.999801 USD

-0.01%

xrp
xrp

$0.993825 USD

-0.72%

solana
solana

$75.624545 USD

0.31%

tron
tron

$0.331431 USD

-0.24%

hyperliquid
hyperliquid

$59.058637 USD

0.40%

dogecoin
dogecoin

$0.069715 USD

-0.57%

unus-sed-leo
unus-sed-leo

$9.445977 USD

0.15%

zcash
zcash

$508.056747 USD

2.72%

monero
monero

$415.302723 USD

-0.23%

chainlink
chainlink

$9.414574 USD

0.05%

cardano
cardano

$0.172880 USD

-2.15%

暗号通貨のニュース記事

NVIDIA GH200 NVL32: リアルタイム AI アプリケーションの最初のトークンまでの時間パフォーマンスを革新

2024/09/27 18:00

NVIDIA の最新の GH200 NVL32 システムは、最初のトークンまでの時間 (TTFT) パフォーマンスが大幅に向上し、Llama 3.1 や 3.2 などの大規模言語モデル (LLM) の増大するニーズに対応します。

NVIDIA GH200 NVL32: リアルタイム AI アプリケーションの最初のトークンまでの時間パフォーマンスを革新

NVIDIA's latest GH200 NVL32 system demonstrates a remarkable leap in time-to-first-token (TTFT) performance, addressing the growing needs of large language models (LLMs) such as Llama 3.1 and 3.2. According to the NVIDIA Technical Blog, this system is set to significantly impact real-time applications like interactive speech bots and coding assistants.

NVIDIA の最新の GH200 NVL32 システムは、最初のトークンまでの時間 (TTFT) パフォーマンスが大幅に向上し、Llama 3.1 や 3.2 などの大規模言語モデル (LLM) の増大するニーズに対応します。 NVIDIA テクニカル ブログによると、このシステムは対話型スピーチ ボットやコーディング アシスタントなどのリアルタイム アプリケーションに大きな影響を与える予定です。

TTFT is the time it takes for an LLM to process a user prompt and begin generating a response. As LLMs grow in complexity, with models like Llama 3.1 now featuring hundreds of billions of parameters, the need for faster TTFT becomes critical. This is particularly true for applications requiring immediate responses, such as AI-driven customer support and digital assistants.

TTFT は、LLM がユーザー プロンプトを処理し、応答の生成を開始するまでにかかる時間です。 Llama 3.1 のようなモデルが数千億のパラメータを備えているため、LLM が複雑になるにつれて、より高速な TTFT の必要性が重要になっています。これは、AI 主導のカスタマー サポートやデジタル アシスタントなど、即時応答が必要なアプリケーションに特に当てはまります。

NVIDIA's GH200 NVL32 system, powered by 32 NVIDIA GH200 Grace Hopper Superchips and connected via the NVLink Switch system, is designed to meet these demands. The system leverages TensorRT-LLM improvements to deliver outstanding TTFT for long-context inference, making it ideal for the latest Llama 3.1 models.

NVIDIA の GH200 NVL32 システムは、32 個の NVIDIA GH200 Grace Hopper スーパーチップを搭載し、NVLink スイッチ システム経由で接続されており、これらの要求を満たすように設計されています。このシステムは TensorRT-LLM の改善を活用して、ロングコンテキスト推論に優れた TTFT を提供し、最新の Llama 3.1 モデルに最適です。

Applications like AI speech bots and digital assistants require TTFT in the range of a few hundred milliseconds to simulate natural, human-like conversations. For instance, a TTFT of half a second is significantly more user-friendly than a TTFT of five seconds. Fast TTFT is particularly crucial for services that rely on up-to-date information, such as agentic workflows that use Retrieval-Augmented Generation (RAG) to enhance LLM prompts with relevant data.

AI スピーチ ボットやデジタル アシスタントなどのアプリケーションでは、人間のような自然な会話をシミュレートするために、数百ミリ秒の範囲の TTFT が必要です。たとえば、0.5 秒の TTFT は、5 秒の TTFT よりもはるかに使いやすいです。高速 TTFT は、検索拡張生成 (RAG) を使用して関連データで LLM プロンプトを強化するエージェント ワークフローなど、最新の情報に依存するサービスにとって特に重要です。

The NVIDIA GH200 NVL32 system achieves the fastest published TTFT for Llama 3.1 models, even with extensive context lengths. This performance is essential for real-time applications that demand quick and accurate responses.

NVIDIA GH200 NVL32 システムは、コンテキストの長さが長い場合でも、Llama 3.1 モデルに対して公開された最速の TTFT を実現します。このパフォーマンスは、迅速かつ正確な応答が要求されるリアルタイム アプリケーションにとって不可欠です。

The GH200 NVL32 system connects 32 NVIDIA GH200 Grace Hopper Superchips, each combining an NVIDIA Grace CPU and an NVIDIA Hopper GPU via NVLink-C2C. This setup allows for high-bandwidth, low-latency communication, essential for minimizing synchronization time and maximizing compute performance. The system delivers up to 127 petaFLOPs of peak FP8 AI compute, significantly reducing TTFT for demanding models with long contexts.

GH200 NVL32 システムは、32 個の NVIDIA GH200 Grace Hopper スーパーチップを接続し、それぞれが NVLink-C2C 経由で NVIDIA Grace CPU と NVIDIA Hopper GPU を組み合わせています。この設定により、同期時間を最小限に抑え、コンピューティング パフォーマンスを最大化するために不可欠な、高帯域幅、低遅延の通信が可能になります。このシステムは、最大 127 ペタフロップスのピーク FP8 AI コンピューティングを提供し、長いコンテキストを伴う要求の厳しいモデルの TTFT を大幅に削減します。

For example, the system can achieve a TTFT of just 472 milliseconds for Llama 3.1 70B with an input sequence length of 32,768 tokens. Even for more complex models like Llama 3.1 405B, the system provides a TTFT of about 1.6 seconds using a 32,768-token input.

たとえば、システムは、入力シーケンス長が 32,768 トークンの Llama 3.1 70B で、わずか 472 ミリ秒の TTFT を達成できます。 Llama 3.1 405B のようなより複雑なモデルの場合でも、システムは 32,768 トークンの入力を使用して約 1.6 秒の TTFT を提供します。

Inference continues to be a hotbed of innovation, with advancements in serving techniques, runtime optimizations, and more. Techniques like in-flight batching, speculative decoding, and FlashAttention are enabling more efficient and cost-effective deployments of powerful AI models.

推論は、サービス提供技術や実行時の最適化などの進歩により、イノベーションの温床であり続けています。実行中のバッチ処理、投機的デコード、FlashAttendant などの技術により、強力な AI モデルのより効率的かつコスト効率の高い導入が可能になります。

NVIDIA's accelerated computing platform, supported by a vast ecosystem of developers and a broad installed base of GPUs, is at the forefront of these innovations. The platform's compatibility with the CUDA programming model and deep engagement with the developer community ensure rapid advancements in AI capabilities.

NVIDIA のアクセラレーテッド コンピューティング プラットフォームは、開発者の広大なエコシステムと GPU の広範なインストール ベースによってサポートされており、これらのイノベーションの最前線にあります。このプラットフォームの CUDA プログラミング モデルとの互換性と開発者コミュニティとの深い関わりにより、AI 機能の急速な進歩が保証されます。

Looking ahead, the NVIDIA Blackwell GB200 NVL72 platform promises even greater advancements. With second-generation Transformer Engine and fifth-generation Tensor Cores, Blackwell delivers up to 20 petaFLOPs of FP4 AI compute, significantly enhancing performance. The platform's fifth-generation NVLink provides 1,800 GB/s of GPU-to-GPU bandwidth, expanding the NVLink domain to 72 GPUs.

将来を見据えると、NVIDIA Blackwell GB200 NVL72 プラットフォームはさらに大きな進歩を約束します。第 2 世代の Transformer Engine と第 5 世代の Tensor コアにより、Blackwell は最大 20 ペタフロップスの FP4 AI コンピューティングを実現し、パフォーマンスを大幅に向上させます。このプラットフォームの第 5 世代 NVLink は、1,800 GB/秒の GPU 間の帯域幅を提供し、NVLink ドメインを 72 GPU に拡張します。

As AI models continue to grow and agentic workflows become more prevalent, the need for high-performance, low-latency computing solutions like the GH200 NVL32 and Blackwell GB200 NVL72 will only increase. NVIDIA's ongoing innovations ensure that the company remains at the forefront of AI and accelerated computing.

AI モデルが成長し続け、エージェント ワークフローがより普及するにつれて、GH200 NVL32 や Blackwell GB200 NVL72 のような高性能で低遅延のコンピューティング ソリューションの必要性は高まる一方です。 NVIDIA は継続的なイノベーションにより、AI とアクセラレーション コンピューティングの最前線にあり続けることが保証されています。

オリジナルソース:blockchain

免責事項:info@kdj.com

提供される情報は取引に関するアドバイスではありません。 kdj.com は、この記事で提供される情報に基づいて行われた投資に対して一切の責任を負いません。暗号通貨は変動性が高いため、十分な調査を行った上で慎重に投資することを強くお勧めします。

このウェブサイトで使用されているコンテンツが著作権を侵害していると思われる場合は、直ちに当社 (info@kdj.com) までご連絡ください。速やかに削除させていただきます。

2026年08月19日 に掲載されたその他の記事