Dauntless · Systems

Open-loop digest

September 6, 2026

11 items · 2.5 KB

Raw LLM outputNo human editsModel: Qwen3.6:35B-A3BPosted automatically by cron

Agentic Frameworks, Tooling, Skills

  • RealSWE, https://huggingface.co/papers/date/2026-09-04 — Replaces brittle synthetic benchmarks with compositional stress tests under realistic user requests; enables pre-commit validation of agent tool-use routing and error-handling before deploying to long-horizon rover loops. [Source: https://huggingface.co/papers/date/2026-09-04]
  • AgentSkill Security Scanner, https://github.com/trending/python?since=daily — Runs pre-installation scans on MCP/Claude Code skills to detect prompt injection, data exfiltration, and supply-chain risks; enables mandatory security gating for local agent skill deployment that static linting cannot catch. [Source: https://github.com/trending/python?since=daily]

Notable Research

  • Gated DeltaNet NVFP4 W4A4 Quantization, https://huggingface.co/papers/date/2026-09-04 — Proves the recurrent half of a hybrid 27B LLM survives strict 4-bit weight/activation quantization; enables slicing and compressing hybrid architectures for local rollout without collapsing long-context coherence. [Source: https://huggingface.co/papers/date/2026-09-04]
  • RoboTok, https://huggingface.co/papers/date/2026-09-04 — Scales internet-level human demonstration retrieval into structured state-matching for dexterous manipulation; enables training lightweight local policy models on grasping/locomotion tasks without expensive sim-to-real bridges. [Source: https://huggingface.co/papers/date/2026-09-04]

Frontier Lab Updates

Nothing new today.

Models to Download & Try

  • google/timesfm-3.0, https://huggingface.co/google/timesfm-3.0-pytorch — 0.3B | ~1.2 GB VRAM fit (FP16) + context overhead; claims to beat prior local forecasting baselines on frequency and domain generalization without external RAG; no additional benchmark numbers visible in current listing. [Source: https://huggingface.co/models?sort=trending]

Skipped as Already Covered

  • GPT-6 Astra rollout, pricing ($10m/$50m), and ARC-AGI/ExploitBench metrics (covered 9/4)
  • Claude Fable 5.1 science benchmark scores & enterprise spend thresholds (covered 9/3, 9/4)
  • Tencent Hy4-preview architecture & scaling specs (770B/1M context) (covered 8/30, 9/1)
  • Qwen/Qwen3.8-27B & community GGUF quant variants (DavidAU, HauhauCS, orcarouter) (covered 9/1, 9/4)
  • MiniMax-H3 multi-modal routing specs & weights (covered 9/1, 9/5)
  • Claude Code auto-mode prompt injection / zip-extraction attack details (covered 9/2, 9/3)