为什么手机内存进入英伟达机柜后,贵过HBM?
过去长期由HBM主导的AI服务器内存账单,正在被“手机内存”改写。 8月11日,外媒Wccftech援引美国银行(BofA)测算称,英伟达面向Vera Rubin Ultra NVL144的Kyber机柜将配置124.4TB HBM4E,成本约为250万美元;同柜216TB LPDDR5X虽然单价更低,但总成本趋近280万美元,超出前者约12%。

AI alone won’t change your business. The system running it will.
(翻译)AI本身不会改变你的业务,运行它的系统才会。
Become an AI-first enterprise with Microsoft’s agent platform.

Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
(翻译)加速科学发现的前沿:Google 对 Genesis Mission 的 4000 万美元承诺
Google is committing $40 million in AI tokens and cloud credits to support the DOE’s Genesis Mission and accelerate groundbreaking scientific discovery.

Up to 3.2x Faster Inference with LFM2.5-DSpark
(翻译)LFM2.5-DSpark 最高可将推理速度提升 3.2 倍
A Blog post by Liquid AI on Hugging Face

Reduce RAG costs on Amazon Bedrock with query-aware compression
(翻译)通过查询感知压缩降低 Amazon Bedrock 上的 RAG 成本
Input tokens are often a meaningful part of the cost of running Retrieval Augmented Generation (RAG) at scale. This post describes a query-aware context compression pattern on Amazon Bedrock: after retrieval, a smaller model filters retrieved chunks against the query before the primary model answers
Introducing cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock
(翻译)在 Amazon Bedrock 上推出 OpenAI GPT-5.6 模型的跨区域推理
Amazon Bedrock now offers OpenAI GPT-5.6 models (Sol, Terra, and Luna) in more than 25 AWS Regions with cross-Region inference. Learn how US geographic and global inference profiles route requests for higher throughput, how to call the models with the OpenAI and Converse APIs, and how to configure I
Building Federated Multimodal AI Workflows with NVIDIA FLARE
(翻译)使用 NVIDIA FLARE 构建联邦多模态 AI 工作流
Modern vision-language models (VLMs) can support tasks such as visual question answering, captioning, and image-text reasoning. In practice, however…

DiffusionGemma: 4x faster text generation
(翻译)DiffusionGemma:文本生成速度最高提升4倍
An overview of DiffusionGemma, an exceptionally fast text generation model with up to 4x faster speeds.

Run Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy
(翻译)使用多个 GPU 在几分钟内运行大规模 UMAP——且不损失精度
Uniform Manifold Approximation and Projection (UMAP) is a dimensionality reduction technique widely used for visualization and feature extraction.

Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer
(翻译)使用 NVIDIA Model Optimizer 通过 QAD 开发 Nemotron 3.5 Lightning NVFP4
Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models…

Granite 4.2 LLMs: How They're Built
(翻译)Granite 4.2 大语言模型:它们是如何构建的
A Blog post by IBM Granite on Hugging Face

Agentic observability with Amazon OpenSearch Service MCP Apps
(翻译)使用 Amazon OpenSearch Service MCP Apps 实现智能体可观测性
Amazon OpenSearch Service now supports MCP Apps, which return interactive visualizations alongside your AI agent's text responses. Learn how a single, locally run MCP server lets your agent move from alert to trace to logs to root cause in one conversation, and how you can verify every step inline w
技术革命:为 Agentic AI 时代做好准备 | 技术趋势
进入 2026 年,随着 Agentic AI 逐渐成熟,智能系统正从被动助手转变为独立行动者。这场革命正在重塑购物体验,也在供应链、门店运营和产品制造等环节深刻改造企业运作方式。

小米低调发布三个芯片,不止瞄准AI手机|最前线
小米发布三款芯片 作者|肖漫 编辑|斯来 无线上直播,雷军也并未到场,仅用了40分钟,小米便发布了三款芯片。 2026 年 8 月 24 日,小米一口气发布了三款玄戒芯片,分别是AI 旗舰 SoC 玄戒O3、 AI 加速芯片玄戒O100,以及智驾高算力 AI 芯片玄戒D100。
