Gemini
23stories this week
33last 30 days
64all time
Timeline
- Oct 10, 2026 · Opinion / analysis · 1 sourceBenefits of using bigger models than Qwen 3.8 flash next?Qwen 3.8 27b was the first model I tried on Ninfer at NVFP4 and then shifted to Flash next after seeing issues with 27b such as not willing to yield to instructions set in AGENTS.md or agent skills.
- Oct 8, 2026 · Research paper · 1 sourceFrom Reactive Containment to Proactive Assurance: Lessons from OpenAI, Anthropic, and Google Agent Security IncidentsIn 2026, cybersecurity evaluations involving OpenAI, Anthropic, and Google agents reached real systems outside their authorized test scope.
- Oct 8, 2026 · Research paper · 1 sourceFastBench: Can Streaming VLMs Perceive High-Dynamic Real-World Streams?Streaming Video Large Language Models (VLMs) enable continuous video understanding, yet existing benchmarks focus on low-dynamic scenarios.
- Oct 8, 2026 · Research paper · 1 sourceIs In-Domain Training Enough for Fine-Grained Industrial Anomaly Understanding?A single multimodal large language model (MLLM) struggles to excel simultaneously at detection, localization, description, and reasoning in multimodal industrial anomaly understanding (MM-IAU).
- Oct 8, 2026 · Research paper · 1 sourceDataVista: Diagnosing Multimodal LLMs on Data Video UnderstandingWe present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains.
- Oct 8, 2026 · Research paper · 1 sourceConstitutional Gating and Deterministic Recovery for Multi-Agent LLM Negotiation: Ablations Against a Stateful Adversarial GatekeeperWe study a three-part control stack - a 5-Pillar runtime constitution, a 4-tier swarm (Director, three-agent majority vote, Monitor, schema hard gate) and Cognitive Annealing (deterministic deadlock detection, atomic purge of the agent-side context, a canonical recovery message) - against a released adversarial Gatekeeper whose acceptance rules are fixed regular expressions and whose LLM only renders reply text.
- Oct 8, 2026 · Research paper · 1 sourceGroundSight at GroundLM 2026 Shared Tasks: GoldenViewVQAWe present CoVeR-VQA, a training-free multi-stage verification and correction framework for grounded multi-view VQA.
- Oct 8, 2026 · Research paper · 1 sourceDeception by Omission: Language Models Knowingly Hide Their MistakesLarge language models (LLMs) increasingly act as agents with little human oversight, so potential mistakes they make can go unnoticed.
- Oct 7, 2026 · Opinion / analysis · 1 sourceDoes better work always mean better workers?But AI is already changing how on-the-job learning works.
- Oct 7, 2026 · Research paper · 1 sourceConversational Voice Aesthetic Model with Reinforcement Learning from Human ListenersWe introduce Conversational Voice Aesthetic Model, a speech large language model for describing the voice aesthetics of real or synthetic speech responses in natural conversational contexts.
- Oct 7, 2026 · Research paper · 1 sourceRECAST: Learning to Compute the Right Context through Adaptive Evidence RoutingIn this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved.
- Oct 7, 2026 · Research paper · 1 sourceDocument-Level Text Simplification in Estonian Using Large Language ModelsDespite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored.
- Oct 7, 2026 · Research paper · 1 sourceFrom Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMsTo bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation.
- Oct 7, 2026 · Research paper · 1 sourceWorldBench: Evaluating LLMs on Three.js Voxel World GenerationWe present WorldBench, a benchmark and judge for open-ended, LLM-generated Three.js worlds.
- Oct 7, 2026 · Research paper · 1 sourceShared and structured inputs undermine collective random choice by reasoning AI agentsRandom selection is widely used in resource allocation and auditing, making reliable implementation essential for AI-agent systems.
- Oct 7, 2026 · Research paper · 1 sourceDual- versus Single-Suggestion AI Support for Radiographic Interpretation in Residents: Randomized Multireader StudyPurpose: To compare dual- and single-suggestion AI support for radiographic interpretation by residents, particularly when the shared AI suggestion was incorrect.
- Oct 7, 2026 · Research paper · 1 sourceArctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic ScienceWe introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence.
- Oct 7, 2026 · Research paper · 1 sourceDo Image Editors Follow Depth-Dependent Blur and Aperture Response? A Rendered-Ground-Truth Pilot AuditA physical aperture edit spreads blur across depth in thin-lens proportions and changes the blur when the aperture changes; prior evaluations check blur monotonicity, sharpness-trend correlation, effective-aperture error, or vision-language judgments, and none we found reports the two properties separately at known depths.
- Oct 6, 2026 · Product / feature launch · 1 sourceEmbeddingGemma 2: an open, lightweight multimodal embedding modelEmbeddingGemma 2: an open, lightweight multimodal embedding model
- Oct 6, 2026 · Research paper · 1 sourceWhen the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using AgentsWe study this possibility in a controlled file-recovery environment where increases in regulatory intensity trigger experimentally imposed tool failures.
- Oct 6, 2026 · Research paper · 1 sourceVerify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving AgentsAs large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML).
- Oct 6, 2026 · Research paper · 1 sourceImage Bitstream Fine-grained Understanding for Privacy-Friendly AIoTImage Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences.
- Oct 5, 2026 · Opinion / analysis · 1 sourcePeople really hate AI, so why can’t they get enough?Over the summer I talked to the CEO of Springboards, a startup building an LLM that’s designed to come up with a wider variety of responses than its mainstream rivals do.
- Oct 3, 2026 · Open-source release · 1 sourcepydantic/pydantic-ai v2.54.0: v2.54.0 (2026-10-02)<!-- Release notes generated using configuration in .github/release.yml at main -->
- Oct 2, 2026 · Open-source release · 1 sourcepydantic/pydantic-ai v2.53.0: v2.53.0 (2026-10-01)This release fixes one security issue in ConcurrencyLimitedModel.
- Sep 30, 2026 · Open-source release · 1 sourcegoogleapis/python-genai v2.26.0
- Sep 30, 2026 · Model release · 1 sourceGemini 4 Argon: our next era of frontier intelligenceGemini 4 Argon: our next era of frontier intelligence
- Sep 30, 2026 · Open-source release · 1 sourcepydantic/pydantic-ai v2.52.0: v2.52.0 (2026-09-29)This release fixes one security issue in webfetch.
- Sep 28, 2026 · Open-source release · 1 sourcecrewAIInc/crewAI 1.15.23Implement evaluation of the last traced run through AMP in crewai eval
- Sep 24, 2026 · Product / feature launch · 1 sourceIntroducing Gemini 3.8 Live with Live AvatarIntroducing Gemini 3.8 Live with Live Avatar