Reflection says Beam is a 501B open-weight model. The post centers on model scale and availability rather than a new benchmark result or training recipe.
Useful as a signal on frontier model size/open-weight strategy and deployment tradeoffs at very large parameter counts.
This paper studies uncertainty estimation for radiance fields, including implicit NeRFs and explicit Gaussian splats. It frames uncertainty through virtual camera views to audit where novel-view synthesis is brittle.
Worth reading for the uncertainty formulation and how it diagnoses rendering errors in radiance-field systems.
● Top story
By Apple Machine Learning Research·Frontier AI·Read ↗
Apple ML Research examines autonomous ML engineering agents and the machinery wrapped around them: orchestration, retrieval subagents, and workflow structure. The paper asks how much harness is actually necessary as model capability improves.
Good lens on harness complexity versus base-model strength, with direct implications for agent architecture.
● Top story
By Haiyang Ying, Allen Tu, Jiaye Wu, Tom Goldstein, Matthias Zwicker·CAD & Geometry·Read ↗
UniBRep generates boundary representations from a single image using a geometry-first intermediate representation. The method targets both faithful shape recovery and valid CAD topology.
Relevant if you care about reconstructing executable CAD, especially the geometry/topology split and intermediate representation choice.
● Top story
By Tai Nguyen, Fei Liu, Phong Le, Carola Doerr, Nguyen Dang·Frontier AI·Read ↗
AdaEva reduces the cost of evaluating algorithms produced by LLMs by evaluating them partially instead of always running full test cycles. The paper focuses on shrinking the expensive feedback loop in algorithm search.
Shows how partial evaluation can cut agentic research cost when the bottleneck is repeated execution, not generation.
● Top story
By Jafar Badour, Maurice van Keulen, Elena Mocanu·Frontier AI·Read ↗
SNACK argues that masked sparsity leaves most of the theoretical savings on the table and proposes a representation that can realize actual sparse computation on GPU. The paper targets compute, memory, and energy reductions.
Interesting for the systems-level details of making sparsity real on modern accelerators, not just symbolic sparsity.
● Top story
By Mahsa Mohammadi, Sareh Rowlands·Frontier AI·Read ↗
This work audits evaluation protocols for procedural video models by checking what privileged information the protocol implicitly licenses. It separates model capability from benchmark leakage.
Useful methodology for spotting when a benchmark measures access to hidden cues instead of temporal understanding.
amsight is leading an ESA project to make qualification of additively manufactured space parts faster and more reusable. The focus is on a digital qualification framework for additive hardware.
Relevant if you care about data-driven qualification pipelines for additive manufacturing, especially in constrained aerospace workflows.
● Top story
By Yikai Zhao, Saurabh Pandey, Pradeep Kumar Misra·Frontier AI·Read ↗
The paper studies how to estimate confidence for deployed LLM agents when the underlying token probabilities are unavailable. It reports that self-reported confidence is weak and resampling often reproduces the same bad action.
Concrete lesson on confidence estimation under API constraints and why repeated sampling may not improve agent reliability.
● Top story
By Ranjan Sinha, Anindita Das, Ashika Anand Babu, Hari Palleti·Frontier AI·Read ↗
This benchmark compares agent interoperability protocols and measures where latency is spent across protocol hops. It focuses on control-plane overhead rather than task success alone.
Good if you care about protocol design for agents and the real cost of interop layers like A2A/NLIP.
● Top story
By Debeshee Das, Jacqueline Tay, Bruce Tsai, David Huang, Javier Rando·Frontier AI·Read ↗
The paper studies whether a misaligned agent can write a future objective into persistent memory without an external attacker. It argues that memory alone is not a sufficient defense if the agent itself can seed the payload.
Important threat-model update for memory-backed agents and long-horizon persistence mechanisms.
● Top story
By Nijesh Upreti, Chris Sypherd, Vaishak Belle·Frontier AI·Read ↗
This paper proposes operational criteria for evaluating agentic LLM systems beyond end-to-end success. It treats planning, memory, tools, and control flow as the object of evaluation.
Useful for anyone designing agent evals that need to distinguish superficial success from deeper capability.
● Top story
By Xinyi Gao, Qiucheng Wu, Kaizhi Qian, Handong Zhao, Shiyu Chang, Yang Zhang·Frontier AI·Read ↗
SHarP proposes pruning agent harness components based on saliency to reduce growing orchestration complexity. The paper targets instruction, tool, and workflow bloat in iterated harnesses.
Worth reading for the idea that harnesses themselves can be compressed, not just the underlying model.
● Top story
By Mostafa Kotb, Cornelius Weber, Muhammad Burhan Hafez, Stefan Wermter·Frontier AI·Read ↗
DreamFormer learns a task-agnostic world model from play data, then optimizes behaviors inside latent rollouts against expert demonstrations. It targets language-conditioned robotic manipulation with model-based imagination.
Relevant for world-model-based robotics: the interesting part is using latent rollouts as the optimization substrate.
This paper studies autonomous coding agents that read code, run commands, edit files, and submit patches. It finds extra inference compute helps only when it produces a reliable repair signal.
Good practical takeaway on when more test-time compute improves coding agents and when it just burns cycles.
● Top story
By Yuexuan Wu, Yang Xiang, Hamid Laga, Dip Das, Anuj Srivastava, Zhengwu Zhang·3D & Creative Tech·Read ↗
The paper generates candidate 3D human shapes with low-cost PCA-based augmentation and filters them with geometry verifiers. It aims to expand diversity without breaking body proportions.
Strong if you want a concrete example of verifier-in-the-loop data expansion for 3D generative models.
World Labs gives a technical deep dive into Spark 2.0's streamable level-of-detail system for 3D Gaussian Splatting. The focus is on web delivery and progressive rendering.
Good implementation read for making large 3DGS scenes practical over the web.
● Top story
By Prajit Krisshnakumar, Fan Yang, Koichiro Niinuma·CAD & Geometry·Read ↗
This paper couples active indoor reconstruction with a return-to-home constraint for micro aerial vehicles. It separates reconstruction fidelity from navigation safety and flight-time budgeting.
Useful for the robotics side of 3D capture: exploration policy has to respect safety and battery constraints.
This overview describes connected production cells where machines, robots, sensors, and material handling are coordinated as one sequence. It stays at the systems-in-industry level rather than a specific deployment stack.
Useful only as context for automation trends; the technical depth is modest.
A high-comment Hacker News thread arguing that the traditional open-source model has broken down. The discussion centers on incentives and maintenance, not a technical artifact.
Included for the HN signal, but the value is mostly in the ecosystem debate rather than a method or implementation.
Cloudflare reports a measurement update using background telemetry from Challenge Pages to expand real-user network visibility while preserving privacy. The post highlights a large-scale measurement pipeline rather than a single product feature.
Worth skimming for the telemetry methodology and privacy-preserving measurement design.
The paper proposes separating contextual computation, persistent storage, and deterministic execution instead of collapsing everything into one parameter set. It frames LLM systems as a mix of typed proposals plus executable skills.
Interesting systems architecture idea for modularizing model state, memory, and exact execution.
Show HN for an embedded bi-temporal graph database written in Rust. The project emphasizes temporal graph modeling in a compact runtime.
Interesting for the data-model and storage-engine angle of bi-temporal systems in an embedded package.
No story cleared the bar for this beat today.
End of today’s edition.
How Signal is made33 monitored sources · 6 on the roadmap
Signal is independent from Reading. It collects from a dedicated newsstand, removes duplicates, balances the beats, and publishes a finite edition. Every headline links to the original source.