AI Intelligence Archive
浏览 AIQB 已保存的 191 stories AI intelligenceArticles,覆盖Models、Products、Industry、Research、教程与观点方法六大分类。
Indexed articles191
Categories6
Selected stories8
All articles
Section 5 / 5 页 · 每页 40 stories
ResearchOrdinary
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-t…
Source: arXiv
ResearchOrdinary
Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecules or through heuristic graph edits that operate directly …
Source: arXiv
ResearchOrdinary
Stochastic Estimation of Transduced Language Models
Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings. Computing the probability of a target prefix under a TLM a…
Source: arXiv
ResearchOrdinary
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both…
Source: arXiv
ResearchOrdinary
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We eval…
Source: arXiv
ResearchOrdinary
Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. No alternative learned direct…
Source: arXiv
ResearchOrdinary
Boosting LLM Exploration via Weak-Model Guidance in RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitiga…
Source: arXiv
ResearchOrdinary
Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. However, deploying message-passing GNNs on a production-scale social graph…
Source: arXiv
ResearchOrdinary
Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organ…
Source: arXiv
ResearchOrdinary
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physi…
Source: arXiv
ResearchOrdinary
How Language Models Organize and Structure Moral Knowledge
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationshi…
Source: arXiv
ResearchOrdinary
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (C…
Source: arXiv
ResearchOrdinary
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymme…
Source: arXiv
ResearchOrdinary
RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature
Retrieved scientific literature can serve as inspiration for both human and AI scientists. Inspiration can take different forms: prior work may directly suggest how to address a problem, or surface directions at different levels of abstrac…
Source: arXiv
ResearchOrdinary
Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions
Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint performance does not establish whether other physiological properties of source contact photoplethysmography (PPG) remain prese…
Source: arXiv
PerspectivesOrdinary
33 Questions Executives Ask About AI-Answered
Source: Every:LatestArticles(网页)
TutorialsOrdinary
Our ChatGPT and OpenClaw Guides Just Got an Overhaul
Source: Every:LatestArticles(网页) · ChatGPT
ModelsOrdinary
Its official now: Ox Alpha is GLM(-5.3 Flash) by zAI. Offical release tonight via zAI Via Bloomberg…
Source: X:Kim (@kimmonismus) · GLM
ProductsOrdinary
Claude for Word: Turn a draft into a finished document
Source: Claude:YouTube(RSS) · Claude
PerspectivesOrdinary
The Case for Cloning Your Coworkers
Source: Every:LatestArticles(网页)
ResearchOrdinary
Ultrafast Frontier Inference: Cerebras Deep Dive at Hot Chips 2026
Source: Cerebras:Blog(网页)
PerspectivesOrdinary
Benchmarks Don't Know Your Job
Source: Every:LatestArticles(网页)
PerspectivesOrdinary
I Tried the AI Model Built to Fix AI Writing
Source: Every:LatestArticles(网页)
IndustrySelected
OpenAI and Anthropic in price war as Chinese AI rivals gain ground
Source: Ars Technica:AI(RSS) · OpenAI / Claude
ModelsSelected
Introducing Fugu-Cyber: our new orchestration model that achieves state-of-the-art performance on real-world cybersecurity benchmarks
Source: Sakana AI:Blog(网页)
ResearchSelected
From Brain Waves to Words: Brain2Qwerty Offers a New Path to Communication Without Surgery
Source: Meta AI:Blog(网页) · Llama
PerspectivesSelected
Sir Demis Hassabis vs Sir Demis Hassabis
Source: Gary Marcus:The Road to AI We Can Trust(RSS)
ModelsSelected
MAI-Image-2.5 launches at No. 2 for image editing on Arena
Source: Microsoft AI:官方博客(网页)
ModelsSelected
Building a hill-climbing machine: Launching seven new MAI models
Source: Microsoft AI:官方博客(网页)
ModelsSelected
DeepSeek-V4Research于Hugging Face发布
DeepSeek-V4 Research已在 Hugging Face 发布 paper: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf
Source: X:AK (@_akhaliq) · DeepSeek / Hugging Face
PerspectivesSelected
i'm joining forces with @ylecun and an incredible group of people to start AMI Labs @amilabs. AMI …
我与 @ylecun 和一群杰出人士联手创立 AMI Labs @amilabs。
Source: X:谢赛宁 (@sainingxie)