

AI 私聊场景的审核错位:为什么”不生成违规内容”被偷换成了”用户违规”
当你把一份包含社区舆情、用户投诉、行业争议事件的竞品调研材料丢给 AI 辅助分析,弹出来的是"内容违规,请遵守社区规范"——这一刻,产品逻辑和场景认知发生了根本错位。本文从一次具体的产品经理工作场景出发,拆解 AI 平台将"公域社区惩戒逻辑"套用到"私域工作场景"的结构性问题,并指出:平台完全有能力在不牺牲合规安全的前提下,做到不影响用户体验。

Cloudflare WriteGuard 为 MCP 服务器提供了精细化的安全控制
Cloudflare 推出了 WriteGuard(目前处于私有测试阶段),旨在为 MCP(模型上下文协议)服务器提供精细化的安全控制。

Introducing Gemini 3.5 Flash Cyber
(翻译)推出 Gemini 3.5 Flash Cyber
Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.
Our approach to bioresilience
(翻译)我们的生物弹性方法
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
微软画图和照片应用生成的图像嵌入了看不见的水印
Solidot是至顶网的科技资讯网站,主要面对开源自由软件和关心科技资讯读者群,包括众多中国开源软件的开发者,爱好者和布道者。口号是“奇客的知识,重要的东西”。
Where Security Fits in an AI Agent Stack
(翻译)安全在AI Agent技术栈中的位置
As AI agents become more capable and operate over longer horizons, building security and trust into the applications they power becomes increasingly important.

Securing the future of AI agents
(翻译)保障AI代理的未来
Securing internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.
Authoring Dogwood policies from natural language in Amazon Bedrock AgentCore
(翻译)在 Amazon Bedrock AgentCore 中通过自然语言编写 Dogwood 策略
AI agents can take actions that do not match your organization's policies. Policy in Amazon Bedrock AgentCore lets teams enforce controls across agents, now including time-based constraints. This post shows how Policy Authoring turns natural-language policy documents into correct Dogwood policies, w
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
(翻译)审计合成回忆录:对照其所描述生活的文档记录度量LLM生成自传中的场景级虚构
Abstract page for arXiv paper 2608.23640: Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
AI Agents Push Humans Out of the Loop
(翻译)AI智能体将人类挤出循环
Abstract page for arXiv paper 2608.23642: AI Agents Push Humans Out of the Loop
Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
(翻译)伦理的LLM辅助研究:负责任委托、验证与认知价值的框架
Abstract page for arXiv paper 2608.23644: Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
(翻译)用于减少医学问答中谄媚与幻觉的门控激活引导
Abstract page for arXiv paper 2608.23666: Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
Automata from Agent Traces: Failure and Next-Step Prediction
(翻译)基于智能体轨迹的自动机:失败与下一步预测
Abstract page for arXiv paper 2608.23670: Automata from Agent Traces: Failure and Next-Step Prediction
A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
(翻译)审计可解释人工智能鲁棒性与保真度的形式化方法论框架:从应用到信任认证
Abstract page for arXiv paper 2608.23817: A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
