AION
Model

Gemini

23stories this week
33last 30 days
64all time

Timeline

  1. Oct 10, 2026 · Opinion / analysis · 1 source
    Benefits of using bigger models than Qwen 3.8 flash next?
    Qwen 3.8 27b was the first model I tried on Ninfer at NVFP4 and then shifted to Flash next after seeing issues with 27b such as not willing to yield to instructions set in AGENTS.md or agent skills.
  2. Oct 8, 2026 · Research paper · 1 source
    From Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security Incidents
    In 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.
  3. Oct 8, 2026 · Research paper · 1 source
    FastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?
    Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
  4. Oct 8, 2026 · Research paper · 1 source
    Is In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?
    A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
  5. Oct 8, 2026 · Research paper · 1 source
    DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
    We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.
  6. Oct 8, 2026 · Research paper · 1 source
    Constitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial Gatekeeper
    We study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text.
  7. Oct 8, 2026 · Research paper · 1 source
    GroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQA
    We present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
  8. Oct 8, 2026 · Research paper · 1 source
    Deception by Omission: Language Models Knowingly Hide Their Mistakes
    Large language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
  9. Oct 7, 2026 · Opinion / analysis · 1 source
    Does better work always mean better workers?
    But AI is already changing how on-the-job learning works.
  10. Oct 7, 2026 · Research paper · 1 source
    Conversational Voice Aesthetic Model with Reinforcement Learning from Human Listeners
    We introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
  11. Oct 7, 2026 · Research paper · 1 source
    RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing
    In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
  12. Oct 7, 2026 · Research paper · 1 source
    Document-Level Text Simplification in Estonian Using Large Language Models
    Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored.
  13. Oct 7, 2026 · Research paper · 1 source
    From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs
    To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation.
  14. Oct 7, 2026 · Research paper · 1 source
    WorldBench: Evaluating LLMs on Three.js Voxel World Generation
    We present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
  15. Oct 7, 2026 · Research paper · 1 source
    Shared and structured inputs undermine collective random choice by reasoning AI agents
    Random selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.
  16. Oct 7, 2026 · Research paper · 1 source
    Dual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader Study
    Purpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
  17. Oct 7, 2026 · Research paper · 1 source
    Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
    We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.
  18. Oct 7, 2026 · Research paper · 1 source
    Do Image Editors Follow Depth-Dependent Blur and Aperture Response? A Rendered-Ground-Truth Pilot Audit
    A physical aperture edit spreads blur across depth in thin-lens proportions and changes the blur when the aperture changes; prior evaluations check blur monotonicity, sharpness-trend correlation, effective-aperture error, or vision-language judgments, and none we found reports the two properties separately at known depths.
  19. Oct 6, 2026 · Product / feature launch · 1 source
    EmbeddingGemma 2: an open, lightweight multimodal embedding model
    EmbeddingGemma 2: an open, lightweight multimodal embedding model
  20. Oct 6, 2026 · Research paper · 1 source
    When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using Agents
    We study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures.
  21. Oct 6, 2026 · Research paper · 1 source
    Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents
    As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
  22. Oct 6, 2026 · Research paper · 1 source
    Image Bitstream Fine-grained Understanding for Privacy-Friendly AIoT
    Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences.
  23. Oct 5, 2026 · Opinion / analysis · 1 source
    People really hate AI, so why can’t they get enough?
    Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do.
  24. Oct 3, 2026 · Open-source release · 1 source
    pydantic/pydantic-ai v2.54.0: v2.54.0 (2026-10-02)
    <!-- Release notes generated using configuration in .github/release.yml at main -->
  25. Oct 2, 2026 · Open-source release · 1 source
    pydantic/pydantic-ai v2.53.0: v2.53.0 (2026-10-01)
    This release fixes one security issue in ConcurrencyLimitedModel.
  26. Sep 30, 2026 · Open-source release · 1 source
    googleapis/python-genai v2.26.0
  27. Sep 30, 2026 · Model release · 1 source
    Gemini 4 Argon: our next era of frontier intelligence
    Gemini 4 Argon: our next era of frontier intelligence
  28. Sep 30, 2026 · Open-source release · 1 source
    pydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)
    This release fixes one security issue in webfetch.
  29. Sep 28, 2026 · Open-source release · 1 source
    crewAIInc/crewAI 1.15.23
    Implement evaluation of the last traced run through AMP in crewai eval
  30. Sep 24, 2026 · Product / feature launch · 1 source
    Introducing Gemini 3.8 Live with Live Avatar
    Introducing Gemini 3.8 Live with Live Avatar

Often appears with