雷峰网发布于 08/24 16:41

为什么手机内存进入英伟达机柜后,贵过HBM?

过去长期由HBM主导的AI服务器内存账单,正在被“手机内存”改写。 8月11日,外媒Wccftech援引美国银行(BofA)测算称,英伟达面向Vera Rubin Ultra NVL144的Kyber机柜将配置124.4TB HBM4E,成本约为250万美元;同柜216TB LPDDR5X虽然单价更低,但总成本趋近280万美元,超出前者约12%。

查看原文
AWS Machine Learning Blog发布于 08/22 00:59

Reduce RAG costs on Amazon Bedrock with query-aware compression

(翻译)通过查询感知压缩降低 Amazon Bedrock 上的 RAG 成本

Input tokens are often a meaningful part of the cost of running Retrieval Augmented Generation (RAG) at scale. This post describes a query-aware context compression pattern on Amazon Bedrock: after retrieval, a smaller model filters retrieved chunks against the query before the primary model answers

查看原文
AWS Machine Learning Blog发布于 08/21 05:46

Introducing cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock

(翻译)在 Amazon Bedrock 上推出 OpenAI GPT-5.6 模型的跨区域推理

Amazon Bedrock now offers OpenAI GPT-5.6 models (Sol, Terra, and Luna) in more than 25 AWS Regions with cross-Region inference. Learn how US geographic and global inference profiles route requests for higher throughput, how to call the models with the OpenAI and Converse APIs, and how to configure I

查看原文
NVIDIA Technical Blog发布于 08/20 01:50

Building Federated Multimodal AI Workflows with NVIDIA FLARE

(翻译)使用 NVIDIA FLARE 构建联邦多模态 AI 工作流

Modern vision-language models (VLMs) can support tasks such as visual question answering, captioning, and image-text reasoning. In practice, however…

查看原文
Google DeepMind Blog发布于 06/11 00:24

DiffusionGemma: 4x faster text generation

(翻译)DiffusionGemma:文本生成速度最高提升4倍

An overview of DiffusionGemma, an exceptionally fast text generation model with up to 4x faster speeds.

查看原文
Hugging Face Blog发布于 08/25 23:14

Granite 4.2 LLMs: How They're Built

(翻译)Granite 4.2 大语言模型:它们是如何构建的

A Blog post by IBM Granite on Hugging Face

查看原文
AWS Machine Learning Blog发布于 08/26 03:00

Agentic observability with Amazon OpenSearch Service MCP Apps

(翻译)使用 Amazon OpenSearch Service MCP Apps 实现智能体可观测性

Amazon OpenSearch Service now supports MCP Apps, which return interactive visualizations alongside your AI agent's text responses. Learn how a single, locally run MCP server lets your agent move from alert to trace to logs to root cause in one conversation, and how you can verify every step inline w

查看原文
36氪发布于 08/26 11:21

小米低调发布三个芯片,不止瞄准AI手机|最前线

小米发布三款芯片 作者|肖漫 编辑|斯来 无线上直播,雷军也并未到场,仅用了40分钟,小米便发布了三款芯片。 2026 年 8 月 24 日,小米一口气发布了三款玄戒芯片,分别是AI 旗舰 SoC 玄戒O3、 AI 加速芯片玄戒O100,以及智驾高算力 AI 芯片玄戒D100。

查看原文