<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel>
<title>雷达站</title>
<link>https://radar.chouzz.com</link>
<description>个人知识雷达：每日采集、去重、归档与解读</description>
<item>
<title>雷达日报 · 2026-09-10</title>
<link>https://radar.chouzz.com/daily/2026-09-10</link>
<guid>https://radar.chouzz.com/daily/2026-09-10</guid>
<pubDate>2026-09-10</pubDate>
<description># 雷达日报 · 2026-09-10

&gt; 今日 65 条信号 · SI 正常 · 精选 4

## 🎯 今日精选

### 1. Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
`2609.09153` · ? · HF 🔥23 · #untagged
**核心 idea**：Large language models are increasingly deployed as agents that plan over long horizons and act through external </description>
</item>
<item>
<title>The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements</title>
<link>https://radar.chouzz.com/items/2509.01809</link>
<guid>https://radar.chouzz.com/items/2509.01809</guid>
<pubDate>2026-09-11</pubDate>
<description>We consider the problem of support recovery for sparse binary signals from noisy linear measurements. For sparse Gaussian measurement matrices we identify sufficient conditions on the minimal sample size for maximum-likelihood recovery in the high-SNR regime ds/p to infty, where p denotes the signal</description>
</item>
<item>
<title>Scaling Automatic Research Agents via World Models</title>
<link>https://radar.chouzz.com/items/2608.12564</link>
<guid>https://radar.chouzz.com/items/2608.12564</guid>
<pubDate>2026-09-11</pubDate>
<description>Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially</description>
</item>
<item>
<title>HyQuant: Hybrid-Precision Quantization for LLM Attention</title>
<link>https://radar.chouzz.com/items/2608.27875</link>
<guid>https://radar.chouzz.com/items/2608.27875</guid>
<pubDate>2026-09-11</pubDate>
<description>Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the attention module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques </description>
</item>
<item>
<title>TempCloze: Can Video-LLMs Identify the Missing Middle?</title>
<link>https://radar.chouzz.com/items/2609.01515</link>
<guid>https://radar.chouzz.com/items/2609.01515</guid>
<pubDate>2026-09-11</pubDate>
<description>Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video</description>
</item>
<item>
<title>From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution</title>
<link>https://radar.chouzz.com/items/2609.02771</link>
<guid>https://radar.chouzz.com/items/2609.02771</guid>
<pubDate>2026-09-11</pubDate>
<description>Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected exampl</description>
</item>
<item>
<title>A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware</title>
<link>https://radar.chouzz.com/items/2609.05463</link>
<guid>https://radar.chouzz.com/items/2609.05463</guid>
<pubDate>2026-09-11</pubDate>
<description>AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local </description>
</item>
<item>
<title>EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents</title>
<link>https://radar.chouzz.com/items/2609.05903</link>
<guid>https://radar.chouzz.com/items/2609.05903</guid>
<pubDate>2026-09-11</pubDate>
<description>Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed</description>
</item>
<item>
<title>OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution</title>
<link>https://radar.chouzz.com/items/2609.06490</link>
<guid>https://radar.chouzz.com/items/2609.06490</guid>
<pubDate>2026-09-11</pubDate>
<description>Recursive Super-Resolution (SR) extends fixed-scale SR to extreme magnification by repeatedly feeding predictions back into the same model, analogous to zooming an image repeatedly. However, ground truth availability at every scale, especially at depth, remains challenging as the required source res</description>
</item>
<item>
<title>PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents</title>
<link>https://radar.chouzz.com/items/2609.06702</link>
<guid>https://radar.chouzz.com/items/2609.06702</guid>
<pubDate>2026-09-11</pubDate>
<description>Sequential memory agents process long documents by reading chunks one after another while maintaining a compact memory state, coupling document traversal to reasoning depth. This coupling introduces sensitivity to evidence placement and ties inference latency linearly to document length. We introduc</description>
</item>
<item>
<title>Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning</title>
<link>https://radar.chouzz.com/items/2609.06806</link>
<guid>https://radar.chouzz.com/items/2609.06806</guid>
<pubDate>2026-09-11</pubDate>
<description>Training strategy, namely whether to retrain from scratch or fine-tune from the previous checkpoint, is an overlooked decision variable in active learning. We show that this choice has exploitable structure: retraining is most useful in early rounds, when each batch can substantially reshape the lab</description>
</item>
<item>
<title>CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation</title>
<link>https://radar.chouzz.com/items/2609.06931</link>
<guid>https://radar.chouzz.com/items/2609.06931</guid>
<pubDate>2026-09-11</pubDate>
<description>Invasive coronary angiography (CAG) is the gold standard for diagnosing coronary artery disease, but interpretation varies substantially among observers. Existing AI systems can improve consistency but lack auditable decision processes and are limited in comprehensive open-ended assessment, undermin</description>
</item>
<item>
<title>SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem</title>
<link>https://radar.chouzz.com/items/2609.07064</link>
<guid>https://radar.chouzz.com/items/2609.07064</guid>
<pubDate>2026-09-11</pubDate>
<description>Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images -- referred to as spatial intelligence -- remains limited. Existing approaches attempt to address this </description>
</item>
<item>
<title>DF26: We Cannot Tell Fake From Real Anymore</title>
<link>https://radar.chouzz.com/items/2609.07369</link>
<guid>https://radar.chouzz.com/items/2609.07369</guid>
<pubDate>2026-09-11</pubDate>
<description>We introduce DF26, a novel benchmark for detecting AI-generated videos containing fully synthetic clips produced by recent text-to-video and image-to-video models. The videos capture single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interv</description>
</item>
<item>
<title>SchemeArena: Factorized Stress Testing of Scheming in LLM Agents</title>
<link>https://radar.chouzz.com/items/2609.08126</link>
<guid>https://radar.chouzz.com/items/2609.08126</guid>
<pubDate>2026-09-11</pubDate>
<description>We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a s</description>
</item>
<item>
<title>SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents</title>
<link>https://radar.chouzz.com/items/2609.08149</link>
<guid>https://radar.chouzz.com/items/2609.08149</guid>
<pubDate>2026-09-11</pubDate>
<description>SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is undermined by two sources of unreliability: reward hacking, enabled by leakage of gold solutions or hidden evalua</description>
</item>
<item>
<title>AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems</title>
<link>https://radar.chouzz.com/items/2609.08572</link>
<guid>https://radar.chouzz.com/items/2609.08572</guid>
<pubDate>2026-09-11</pubDate>
<description>Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language fee</description>
</item>
<item>
<title>PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving</title>
<link>https://radar.chouzz.com/items/2609.08965</link>
<guid>https://radar.chouzz.com/items/2609.08965</guid>
<pubDate>2026-09-11</pubDate>
<description>Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysi</description>
</item>
<item>
<title>SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators</title>
<link>https://radar.chouzz.com/items/2609.09155</link>
<guid>https://radar.chouzz.com/items/2609.09155</guid>
<pubDate>2026-09-11</pubDate>
<description>World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: chan</description>
</item>
<item>
<title>StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean</title>
<link>https://radar.chouzz.com/items/2609.09264</link>
<guid>https://radar.chouzz.com/items/2609.09264</guid>
<pubDate>2026-09-11</pubDate>
<description>Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at vary</description>
</item>
<item>
<title>MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes</title>
<link>https://radar.chouzz.com/items/2609.10016</link>
<guid>https://radar.chouzz.com/items/2609.10016</guid>
<pubDate>2026-09-11</pubDate>
<description>We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In</description>
</item>
<item>
<title>Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States</title>
<link>https://radar.chouzz.com/items/2609.10060</link>
<guid>https://radar.chouzz.com/items/2609.10060</guid>
<pubDate>2026-09-11</pubDate>
<description>Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants,</description>
</item>
<item>
<title>The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding</title>
<link>https://radar.chouzz.com/items/2609.10296</link>
<guid>https://radar.chouzz.com/items/2609.10296</guid>
<pubDate>2026-09-11</pubDate>
<description>Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical re</description>
</item>
<item>
<title>Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs</title>
<link>https://radar.chouzz.com/items/2609.10355</link>
<guid>https://radar.chouzz.com/items/2609.10355</guid>
<pubDate>2026-09-11</pubDate>
<description>Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal gro</description>
</item>
<item>
<title>Optimizing AI Inference Across the Deployment Stack</title>
<link>https://radar.chouzz.com/items/2609.10550</link>
<guid>https://radar.chouzz.com/items/2609.10550</guid>
<pubDate>2026-09-11</pubDate>
<description>AI deployment performance is shaped not by model architecture alone, but by interactions among compression, compiler transformations, and serving policies. Published benchmarks often report latency and throughput under incomparable conditions, limiting their use for deployment decisions. This paper </description>
</item>
<item>
<title>Some Early Results by Tutte Regarding the Cycle Double Cover Conjecture in 1948</title>
<link>https://radar.chouzz.com/items/2609.10618</link>
<guid>https://radar.chouzz.com/items/2609.10618</guid>
<pubDate>2026-09-11</pubDate>
<description>OpenAI recently announced a proof of the Cycle Double Cover (CDC) Conjecture. Most media reports have characterized it as a 50-year-old open problem. In reality, according to a 1987 letter from Tutte to Fleischner, the Cycle Double Cover Problem has been open for at least 80 years. Two early results</description>
</item>
<item>
<title>Energy deposition in planetary and exoplanetary atmospheres induced by cosmic rays</title>
<link>https://radar.chouzz.com/items/2609.10688</link>
<guid>https://radar.chouzz.com/items/2609.10688</guid>
<pubDate>2026-09-11</pubDate>
<description>Cosmic rays can significantly alter the abundances of certain species in the upper layers of planetary atmospheres, especially in terms of their biosignatures. To fully understand the extent of this effect, it is essential to accurately model the interactions of cosmic rays with planetary magnetic f</description>
</item>
<item>
<title>An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics</title>
<link>https://radar.chouzz.com/items/2609.10712</link>
<guid>https://radar.chouzz.com/items/2609.10712</guid>
<pubDate>2026-09-11</pubDate>
<description>We study how model post-training and test-time inference design affect natural-language proof generation for hard olympiad mathematics. Starting from Nemotron 3 Ultra, we train two specialist checkpoints using supervised fine-tuning and reinforcement learning, and evaluate checkpoint choice, verific</description>
</item>
<item>
<title>NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction</title>
<link>https://radar.chouzz.com/items/2609.10715</link>
<guid>https://radar.chouzz.com/items/2609.10715</guid>
<pubDate>2026-09-11</pubDate>
<description>We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and mo</description>
</item>
<item>
<title>Alternative AI Philosophy: Daoism as Method for AI in Education</title>
<link>https://radar.chouzz.com/items/2609.10842</link>
<guid>https://radar.chouzz.com/items/2609.10842</guid>
<pubDate>2026-09-11</pubDate>
<description>As artificial intelligence (AI) rapidly iterates and transforms teaching, learning, and knowledge production, philosophical reflection has become increasingly indispensable to educational debates that remain predominantly shaped by Western intellectual traditions. This article proposes Daoism as an </description>
</item>
<item>
<title>Learned Continuous Synthesis of Quadratic Difference Tone Spectra</title>
<link>https://radar.chouzz.com/items/2609.10913</link>
<guid>https://radar.chouzz.com/items/2609.10913</guid>
<pubDate>2026-09-11</pubDate>
<description>Quadratic difference tones (QDTs) are a species of auditory distortion product in which a "phantom" pure tone, absent from the acoustic signal, is clearly audible to listeners. Exploiting this phenomenon, one can synthesize harmonically rich tones for musical purposes, a technique called Quadratic D</description>
</item>
</channel></rss>