InfoQ 中文发布于 08/23 17:10

事故频发并不意味着可靠性下降

IT 工程项目领导者最常见的假设之一,是报告的事故数量不断增加就意味着系统可靠性正在下降。然而,Great Circle 最近的一篇文章提出,实际情况往往恰恰相反:事故数量增加,反而可能意味着组织的事故管理文化正在改善。随着团队不断投入资源完善流程、工具、培训和运维规范,他们会更愿意正式地将那些过去大概只会被悄悄处理、甚至被隐瞒的事故报告上去。这样带来的结果就是组织对运营问题的可见性提高了,但却不一定意味着系统的健康状况恶化。

查看原文
InfoQ 中文发布于 08/23 01:17

Cloudflare 推出 Cache Response Rules,在源站响应后进一步控制缓存

Cloudflare 近日推出缓存响应规则(Cache Response Rules),这是一套新的规则引擎,运行在源站返回响应之后、内容写入 Cloudflare 缓存之前。此前,缓存规则(Cache Rules)只能根据请求属性进行判断;缓存响应规则则新增了一个响应处理阶段,可以在响应进入缓存之前检查源站返回的内容。

查看原文
Microsoft AI Blog发布于 07/21 02:20

Microsoft expands Azure AI and HPC infrastructure with AMD

(翻译)微软借助AMD扩展Azure AI和HPC基础设施

AI workloads are scaling faster than any single infrastructure approach can support — with more models, new agent-driven workloads and surging compute demand driving the need for greater specialization across the stack. To meet this need, Microsoft continues to evolve Azure’s infrastructure, includi

查看原文
NVIDIA Technical Blog发布于 08/24 23:00

Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS

(翻译)借助 NVIDIA DSX MaxLPS 最大化 AI 工厂每瓦性能

AI factories are power-constrained industrial systems. The question is no longer how many GPUs fit in a data center, but how much AI output each available…

查看原文
AWS Machine Learning Blog发布于 08/26 00:35

Governed reports with Amazon Quick Desktop and Amazon FSx for NetApp ONTAP

(翻译)使用Amazon Quick Desktop和Amazon FSx for NetApp ONTAP实现受治理的报告

Build a governed weekly reporting workflow with Amazon Quick Desktop and Amazon FSx for NetApp ONTAP. An Amazon S3 access point exposes an approved folder to a Quick knowledge base, and a custom skill drafts cited weekly reports and Slack summaries with human review before anything is shared.

查看原文
AWS Machine Learning Blog发布于 08/26 03:00

Agentic observability with Amazon OpenSearch Service MCP Apps

(翻译)使用 Amazon OpenSearch Service MCP Apps 实现智能体可观测性

Amazon OpenSearch Service now supports MCP Apps, which return interactive visualizations alongside your AI agent's text responses. Learn how a single, locally run MCP server lets your agent move from alert to trace to logs to root cause in one conversation, and how you can verify every step inline w

查看原文