AI News Digest 2026-09-24
特集
中堅コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧
参考記事一覧を表示
- Claude Opus 5.5 (importance:85 / dev:75)
- OpenAI's GPT-6 Sol and Luna cut prices in half but barely move the needle on performance (importance:75 / dev:80)
- Gemini 3.8 text-to-speech (importance:60 / dev:70)
- Anthropic’s biolab made a discovery it’s comparing to Crispr (importance:75 / dev:30)
- Radicle: Disclosure of Vulnerability in the Network Protocol (importance:65 / dev:85)
- OpenAI extends cyber access to Ukraine for civilian defense (importance:20 / dev:10)
- MiMo-V2.6 (importance:60 / dev:70)
- 「i-have-adhd」を日本語で評価したら、公式より良い結果が出た話 (importance:0 / dev:75) (いいね相当スコア: 0)
- Chinese AI Firms Fall on Report of DeepSeek, Moonshot Probe - Bloomberg (importance:30 / dev:15)
- DeepSeek to brief UN Security Council on AI risks amid rising global tech rivalry - The News International (importance:35 / dev:25)
- DeepSeek publishes paper on agent training system DSec with Liang Wenfeng - kucoin.com (importance:60 / dev:70)
- DeepSeek to use Huawei chips widely for AI training over Nvidia - Huawei Central (importance:50 / dev:40)
- Chinese hacker deployed Anthropic and Deepseek and Moonshot AI agents in massive 100-company cyberattack - oodaloop.com (importance:70 / dev:75)
- SpaceXAI Launches Grok 4.7 With Improved Coding, Knowledge Work And AI Safety Capabilities - Pulse 2.0 (importance:70 / dev:75)
- About a month after its launch, Grok Bot, the AI agent under SpaceX (SPCX.US), surpassed 400,000 weekly active users. - Moomoo (importance:40 / dev:30)
- Grok in your Tesla just got connected to your inbox, calendar, and files - pocnetwork.net (importance:55 / dev:50)
- Elon Musk Just Said "Dario Is Right" About Slowing AI Development. Here's What That Endorsement Could Mean for Tesla's and xAI's AI Road Maps. - The Globe and Mail (importance:30 / dev:20)
- Fixing the Portobello Police Station Clock (importance:0 / dev:0)
- A brief history of Windows scroll bar shortcuts (importance:10 / dev:50)
- Jev in 25 Lines of Python (importance:50 / dev:75)
- Claude Code reads AGENTS.md only when telemetry is on [fixed] (importance:40 / dev:75)
- 28% of job postings on company career sites have been open over 90 days (importance:15 / dev:50)
- Stripe's Knowledge AI Platform (importance:65 / dev:70)
- Seattle City Council votes to ban surveillance pricing in sale of groceries (importance:15 / dev:10)
- I don't want the details (importance:15 / dev:45)
- Strands Harness (importance:70 / dev:85)
- Tokens Too Cheap to Meter (importance:65 / dev:80)
- Web-based IBM 1620 emulator and IPL-V from 1963 (importance:10 / dev:60)
- QuestDB (YC S20) Is Hiring a Sales Engineer (importance:5 / dev:30)
- OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005 (importance:70 / dev:50)
- GPT-6 Astra has gained the ability to drive a car (importance:60 / dev:40)
- Transit rewards (importance:15 / dev:10)
- What California is learning from solar panels built over irrigation canals (importance:10 / dev:0)
- The GitHub wiki is an anti-pattern (2022) (importance:30 / dev:75)
- Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived (importance:35 / dev:60)
- SAML: A fractal of bad design (importance:40 / dev:75)
- 'We hacked the FBI:' Hackers say they have data on all FBI employees (importance:60 / dev:40)
- The Situation Report (importance:45 / dev:40)
- v2.1.280 (importance:55 / dev:85)
- Two years of OpenAI Academy (importance:25 / dev:50)
- How invideo improves color grading 3x with GPT‑6 Astra (importance:40 / dev:30)
- Ringg’s AI agents resolve up to 65% of customer calls with OpenAI (importance:45 / dev:45)
- Harvey turns legal context into stronger drafts with GPT-6 Astra (importance:40 / dev:30)
- Grab and OpenAI bring practical AI skills to Southeast Asia (importance:40 / dev:60)
- Better prompt caching for GPT-6 (importance:65 / dev:85)
- Parallel cut research time and cost in half with GPT‑6 Astra (importance:50 / dev:50)
- Priorities and principles for effective third party assessments (importance:50 / dev:60)
- Google Beam expands with new regions, partners, and customers (importance:45 / dev:50)
- Advancing Private AI Compute with secure, server-side memory (importance:55 / dev:75)
- Offloaded inference for real-world physical AI robotics (importance:55 / dev:65)
- How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows (importance:35 / dev:70)
- How UK AISI and EvalEval Are Making Benchmark Results Reproducible (importance:50 / dev:65)
- Transformers now runs llama.cpp quants (importance:55 / dev:80)
- Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community (importance:30 / dev:60)
- Rendering huge pull requests in the GitHub Copilot app (importance:55 / dev:85)
- Developers want more efficient software. Here’s what over 1000 GitHub users told us they need. (importance:50 / dev:80)
- We just shipped support for the ugliest part of HTTP: Vary (importance:35 / dev:80)
- Introducing Worker Previews: Isolated preview environments for every change your agent makes (importance:55 / dev:85)
- Small Talk With Prasun Kumar, CEO and Founder of Oppex AI (importance:20 / dev:30)
- 100 Exercises to Learn Rust, Updated (importance:40 / dev:75)
- Code Quality Q&A With the JetBrains Qodana Team (importance:40 / dev:80)
- JetBrains Air: Building a System of Products for Agentic Software Development (importance:70 / dev:85)
- Meet the Ecosystem: Partners and Customers at WeAreDevelopers with Docker (importance:25 / dev:60)
- Manage Kubernetes Node Fleets with NodeWright (importance:45 / dev:80)
- How SWE-Serve Exposes the Gap Between Local Tests and Live Serving (importance:50 / dev:85)
- Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing (importance:60 / dev:80)
- Topology-Aware Workload Scheduling with NVIDIA Topograph (importance:50 / dev:80)
- What’s New for Game Developers: DLSS 5 with 3D-Guided Neural Rendering, NVIDIA ACE Updates, and New RTX Kit Capabilities (importance:55 / dev:75)
- Roundtables: The Deadly Failures of The Virtual Border Wall (importance:15 / dev:15)
- Don’t be fooled by this summer of AI hype (importance:50 / dev:60)
- Owners mourn spoiled food after firmware update bricks Samsung smart fridges (importance:15 / dev:40)
- New Anthropic, OpenAI models make same promise: A little more for a lot less money (importance:65 / dev:80)
- Microsoft disrupts AI-assisted platform that compromised 12,000 accounts (importance:55 / dev:70)
- Lawsuit demands OpenAI pay for new school after ChatGPT used in shooting (importance:35 / dev:20)
- Toyota orders workers to train humanoid robots but says humans won't be replaced (importance:40 / dev:40)
- Dyson’s most overengineered gadget may have a waterproofing problem (importance:10 / dev:20)
- ChatGPT mobile app gets voice-based agentic features (importance:55 / dev:50)
- Even Americans who use AI every day are worried about it (importance:30 / dev:40)
- YouTube Music gets more conversational with new AI features (importance:35 / dev:30)
- YouTube will let you build your own algorithm with AI (importance:50 / dev:40)
- StrictlyVC at TechCrunch Disrupt 2026: Inside the changing rules of venture capital (importance:15 / dev:20)
- YouTube releases new AI features for creators within its Studio app (importance:45 / dev:50)
- 3 days left to save up to $200 and make impactful connections at TechCrunch Disrupt 2026 (importance:5 / dev:10)
- Spotify is giving you the keys to its recommendation algorithm with US launch of ‘Taste Profile’ (importance:45 / dev:50)
- Ema raises $77M as AI starts eating into enterprise software and services (importance:50 / dev:60)
- ‘We’re already fighting yesterday’s battle’: Greece’s prime minister gets candid about AI (importance:25 / dev:30)
- TechCrunch Founder Summit’s agenda revealed: Unlock fundraising, hiring, and AI insights in Boston on November 4 (importance:15 / dev:30)
- Snorkel AI triples valuation to $3.5B as demand for AI training data booms (importance:55 / dev:70)
- Qualcomm launches two new smartphone chips with emphasis on AI (importance:65 / dev:75)
- Meta admits Muse’s likeness to OpenClaw isn’t a coincidence (importance:40 / dev:40)
- AstroForge is putting AI in command of its next spacecraft (importance:55 / dev:50)
- Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda (importance:35 / dev:60)
- TechCrunch Disrupt 2026: Aaron Edsinger brings Hello Robot’s Stretch 4 to life onstage (importance:30 / dev:40)
- Exhibit tables added: One last chance to showcase your startup at TechCrunch Disrupt 2026 (importance:5 / dev:10)
- Meta’s AI agent is a cute little guy who’s great at spending my money (importance:35 / dev:40)
- Data centers are black boxes, but California wants to change that (importance:40 / dev:50)
- Bernie Sanders proposes banning ‘superintelligence’ and putting violators in prison (importance:40 / dev:20)
- YouTube is building AI creator tools that do almost everything for them (importance:50 / dev:40)
- OpenAI nabs key Patreon execs ahead of upcoming announcement (importance:40 / dev:50)
- OpenAI wants to consult elite mathematicians about how to not fumble again (importance:45 / dev:50)
- Rabbit’s new AI agent doesn’t need an R1 to run (importance:48 / dev:38)
- Andreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta (importance:30 / dev:25)
- Graphify: Unifying Codebase Context to Streamline Agentic Software Engineering (importance:68 / dev:82)
- Beyond Kubernetes at Modal: How to Scale 1 Million Concurrent Sandboxes in Seconds (importance:65 / dev:88)
- Presentation: APIs for Agents: Rethinking API Programs in the MCP Era (importance:70 / dev:88)
- Google Adds Cycle-Level Kernel Profiling to XProf (importance:45 / dev:80)
- Google Open-Sources AX a Kubernetes Style Orchestrator for Autonomous AI Agents (importance:75 / dev:88)
- GitLab Duo Expands Self-Hosted AI Options Through Microsoft Foundry (importance:55 / dev:75)
- Why Read a Research Paper When You Can Turn It Into an AI Agent? (importance:50 / dev:55)
- The Future Is Fanless: 100% Heat Capture for Liquid Cooled AI Servers (importance:38 / dev:35)
- ChatGPT Voice gets closer to "Her" with email, calendar, and Slack access (importance:55 / dev:45)
- Google's new Flash TTS models let you design AI voices from scratch using text descriptions (importance:55 / dev:65)
- YouTube adds AI tools to Creator Studio with script coaching, smart thumbnails, and Gemini editing (importance:45 / dev:45)
- Anthropic engineer explains why Claude's writing got worse although the model got smarter (importance:50 / dev:65)
- Nvidia-backed Nscale keeps its biggest customer, Bytedance, out of its IPO filing (importance:35 / dev:30)
- Meta's AI agent Muse draws 500,000 users in a week along with claims it copied OpenClaw (importance:55 / dev:45)
- Inside Basecamp Research, the AI startup turning evolution into training data (importance:60 / dev:45)
- Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent (importance:60 / dev:75)
- OpenAI hires Patreon co-founder Sam Yam to lead a new Creator Product division (importance:35 / dev:20)
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation (importance:40 / dev:35)
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN Optimization (importance:40 / dev:55)
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian Splatting (importance:45 / dev:65)
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival Prediction (importance:40 / dev:55)
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings (importance:60 / dev:75)
- X-Planner: Event-Structured Task Planning for Embodied Intelligence (importance:55 / dev:75)
- Lean Pool: An AI-Maintained Archive of Formalized Mathematics (importance:60 / dev:75)
- The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis (importance:50 / dev:65)
- MedGate-Fusion: Integrating First-Encounter Semantic Narratives and Physiological Biomarkers for Prospective Stroke Risk Stratification (importance:45 / dev:55)
- When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning (importance:55 / dev:75)
- Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass (importance:60 / dev:75)
- Learned Enterprise Data Comprehension: Compression and Routing for Data Agents (importance:65 / dev:82)
- Making Agents More Consistent: Skills Should Form Habits for Repeat Tasks (importance:65 / dev:82)
- Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing (importance:40 / dev:55)
- Clarification Is Not Correction: LLMs Fail to Let Go (importance:55 / dev:75)
- From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought (importance:60 / dev:80)
- Robust Failure, Conservative Repair: Textual Knowledge Distillation from Cross-Model Failures (importance:55 / dev:80)
- Efficient Iterative Retrieval with Heterogeneous Batching (importance:60 / dev:82)
- From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI (importance:55 / dev:75)
- ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent Runtimes (importance:70 / dev:88)
- Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions (importance:65 / dev:82)
- Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine Learning (importance:50 / dev:70)
- ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations (importance:60 / dev:80)
- RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation (importance:65 / dev:82)
- Spectra: A Rules-Driven LLM Pipeline for Automated KYC Document Processing (importance:60 / dev:80)
- Queer inclusion in speech datasets: An audit and taxonomy of practical tensions (importance:50 / dev:55)
- Towards participatory speech dataset curation: A queer case study and conceptual framework (importance:50 / dev:60)
- SMTB: Fast Structure-Mapping with Tight Bounds (importance:45 / dev:65)
- Weakly Supervised Quantum Error Mitigation (importance:55 / dev:75)
- Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding (importance:60 / dev:80)
- A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators (importance:55 / dev:75)
- Direct Optimization of Generators for Search in Automated Theorem Proving (importance:60 / dev:80)
- Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study (importance:45 / dev:55)
- Transformer Heads Looking for Order (importance:55 / dev:75)
- Evaluating Coding Agents on Kernel Exploit Generation (importance:70 / dev:88)
- ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications (importance:60 / dev:80)
- Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA (importance:60 / dev:82)
- ChatT2: An Adaptive Framework for Developing a Large Language Model-Based Agent for Natural Product Domain Research (importance:50 / dev:65)
- Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces (importance:60 / dev:82)
- Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark (importance:60 / dev:82)
- Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing (importance:55 / dev:70)
- Toolcompass: Guiding Tool Trialing, Not Suppressing It (importance:65 / dev:82)
- How Strongly Should Task State Influence an LLM Agent? (importance:65 / dev:82)
- TCMaster: Confidence-Aware Querying and Workload-Guided Physical Design for Multi-Source Traditional Chinese Medicine Knowledge Graphs (importance:50 / dev:65)
- LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data (importance:50 / dev:65)
- OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities (importance:60 / dev:80)
- The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance (importance:60 / dev:75)
- Neurosymbolic Action Model Learning under Partial Observability (importance:60 / dev:82)
- Towards Omni-dimensional GUI Agent Navigation with Masked Trajectory Prediction (importance:65 / dev:88)
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks (importance:65 / dev:82)
- When Are Aggregate Agent Traces Diagnosable? Traffic-Governed Interpretation and Calibrated Abstention (importance:65 / dev:88)
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts (importance:65 / dev:82)
- Prediction Is Not Detection: Evaluating Pre-Recognition Claims in Longitudinal Clinical AI (importance:55 / dev:75)
- AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing (importance:55 / dev:80)
- CausalLoss-Fin: Attributing Financial-Agent Loss to Decisions and Infrastructure Faults (importance:60 / dev:82)
- VideoX-Qwen: Data-Centric Instruction-Based Video Editing (importance:60 / dev:75)
- CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions (importance:60 / dev:82)
- Canonical locks that encode part-whole hierarchies (importance:55 / dev:80)
- FIRE: Failure-Informed Runtime Engineering for Reliable Language-Model Agents (importance:70 / dev:88)
- Adversarial Course-of-Action Generation: Game-Theoretic Multi-Agent Algorithms for COA matching & COA generation (importance:60 / dev:82)
- ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models (importance:65 / dev:82)
- RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty (importance:55 / dev:80)
- Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows (importance:65 / dev:88)
- The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale (importance:60 / dev:80)
- FusionMMT: A Unified Multimodal and Multitask Learning Framework for Nuclear Fusion (importance:55 / dev:75)
- Neoadjuvant chemotherapy response prediction using pretreatment diffusion and contrast-enhanced magnetic resonance imaging with clinical variables (importance:50 / dev:65)
- Early Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy Using Temporal Deep Learning on DWI (importance:50 / dev:65)
- DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents (importance:70 / dev:88)
- FairMon: A Tool for Monitoring and Visualizing Algorithmic Fairness (importance:65 / dev:82)
- MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation (importance:55 / dev:75)
- When Big Data Becomes a Curse: Spatial Heterogeneity and the Limits of Learning from Passive Acoustic Monitoring Data (importance:45 / dev:65)
- The Cost of Conservation: Coordination-Memory Laws for Exact-Support Generation (importance:60 / dev:82)
- VACS: Value-Aligned Compositional Shielding for Multi-Agent Reasoning (importance:65 / dev:88)
- When Verifiers Vote Backwards under Verdict Substitution: Signed Pivotal Value in Correlated Self-Consistency (importance:60 / dev:80)
- Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs (importance:60 / dev:80)
- Toward User-Mediated Self-Repair in Ubiquitous Robots Through Goal-Oriented Agentic AI (importance:60 / dev:80)
- The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke (importance:60 / dev:82)
- EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models (importance:65 / dev:82)
- TREND-10K: A Comprehensive Dataset for Next-Generation Video Quality Assessment Based on Preference-Driven Media (importance:50 / dev:65)
- Identifying Intelligent Processes via Online Sequential Testing (importance:60 / dev:80)
- RCShift: Certifying When Partial Linkage Suffices for Finite-Sample Decisions (importance:55 / dev:80)
- Improved Multiplayer Bandit Algorithm for Bernoulli Rewards (importance:40 / dev:20)
- A Hybrid AI Framework for Academic Advising: Integrating Ensemble-Based Grade Prediction and a Rule-Based Expert System (importance:30 / dev:45)
- A Multi-Timestep LSTM Ensemble regressor for Enhanced Short-Term Runoff Prediction (importance:30 / dev:35)
- Coding Agents are Strong Prompt Optimizers (importance:60 / dev:75)
- Decoupling Is Not Identification: Supervised Evidential Learning in Next-Token Prediction (importance:50 / dev:55)
- FISSION: Label Augmentation for Bot Detection (importance:50 / dev:60)
- Dual-Frontier: When Can an Agent Trust Its World Model? (importance:55 / dev:75)
- Reliability Theory for AI Control (importance:70 / dev:80)
- The Source of Disturbance Matters: External, Internal, and Control-Generated Noise in Adaptive Regulation (importance:35 / dev:45)
- Recursive self-improvement of AI research agents (importance:75 / dev:85)
- Reproducible AI Requires Reproducible Randomness (importance:60 / dev:75)
- REFLEX with Jev for Efficient Selective Control in LLM Agents (importance:65 / dev:80)
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure (importance:60 / dev:75)
- Neutral-Atom-based Quantum Optimization for Resource Allocation in NOMA Networks (importance:50 / dev:55)
- Quantum-Aided Active Device Detection in Energy-Harvesting Symbiotic Radio Networks (importance:45 / dev:55)
- The Delegation Blind Spot: Auditing Product Decisions from Agent Choices (importance:60 / dev:80)
- Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It (importance:60 / dev:80)
- Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents (importance:70 / dev:85)
- SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving (importance:75 / dev:85)
- CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents (importance:70 / dev:85)
- Financial sentiment analysis using FinBERT with application in predicting stock movement (importance:40 / dev:55)
- Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models (importance:65 / dev:80)
- "As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It (importance:55 / dev:75)
- AIBuildAI-2.5: Efficient Autonomous AI Model Development Through LLM-Guided Tree Search (importance:75 / dev:85)
- Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation (importance:60 / dev:75)
- Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione (importance:65 / dev:80)
- FrontierMath Erd\H{o}s (importance:65 / dev:65)
- LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping (importance:40 / dev:55)
- Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures (importance:60 / dev:80)
- LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay (importance:70 / dev:85)
- Physics-guided deep metric learning with continuous time embeddings for open-world radar pulse de-interleaving (importance:40 / dev:55)
- Federating Quantum and Classical Computing: A Privacy-Preserving Hybrid Approach (importance:50 / dev:65)
- You've Seen Enough: Quality-Constrained Image Coding for Machines (importance:55 / dev:65)
- Rachel: A general-purpose language model directs and revises retrosynthetic routes (importance:55 / dev:55)
- WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation (importance:60 / dev:55)
- Impact Is Not Invalidation: Ask About the Claim, Not the Diff (importance:60 / dev:85)
- The Probabilistic Structure of Large Language Models (importance:65 / dev:75)
- Stable Unsupervised Continual Chunking with Sheaf SyncMap (importance:50 / dev:65)
- Brain-Inspired Hierarchical Modularity for General Continual Learning (importance:55 / dev:70)
- Towards Sustainable Magnetic Resonance Imaging: Insights from long-term, high-resolution energy recordings across an entire scanner fleet (importance:40 / dev:30)
- Exposing Blind Spots in Deep Imbalanced Regression Evaluation (importance:55 / dev:75)
- Benchmarking Neural Defend ARCAS 1B: A Foundational Multimodal Deepfake Detection Model (importance:60 / dev:70)
- Mitigating Sequential Reappearance in Diffusion Data-Point Unlearning (importance:55 / dev:70)
- Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction (importance:75 / dev:80)
- Multi-Term Fourier Graph Neural Network with Sample Relationship Learning for Enhanced Remaining Useful Life Prediction (importance:50 / dev:65)
- From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health (importance:55 / dev:60)
- GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings (importance:60 / dev:75)
- Indirect tipping: a social attack surface in AI agent populations (importance:70 / dev:85)
- Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers (importance:70 / dev:85)
- How Children Design and Reason about Trustworthy AI Chatbots (importance:40 / dev:55)
- Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes (importance:55 / dev:70)
- VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models (importance:65 / dev:80)
- PICPIs: Prediction-Interval-Conditional Prediction Intervals (importance:50 / dev:65)
- Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development (importance:75 / dev:90)
- Deep Reinforcement Learning on Item-Compatibility Graphs for One-Dimensional Bin Packing (importance:55 / dev:75)
- Beyond Natural Language: An Agent-Native Language for Autonomous Science (importance:80 / dev:90)
- Predictive Uncertainty for Neural CAE Surrogates (importance:55 / dev:70)
- Lightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems (importance:65 / dev:85)
- Transformer-Informed Trajectory Optimization for Relative Motion in Cislunar Orbits (importance:55 / dev:65)
- A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization (importance:65 / dev:80)
- Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining (importance:65 / dev:80)
- RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models (importance:70 / dev:85)
- Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains (importance:65 / dev:85)
- Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training (importance:70 / dev:85)
- West-WRF AI 2-km: High-Resolution Prediction of Integrated Vapor Transport and Precipitation (importance:55 / dev:55)
- Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference (importance:75 / dev:90)
- A JEPA Recipe for Tabular Foundation Models (importance:60 / dev:75)
- DefaultGNN: A Dual-Perspective GNN Framework for Predicting Corporate Default from Buyer-Seller Transaction Networks (importance:55 / dev:65)
- IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models (importance:70 / dev:85)
- AkasicMEM: Governed Enterprise Memory for Agents (importance:75 / dev:90)
- RootQuantV2: Adapting a Vision Foundation Model for Root-Trait Regression from Minirhizotron Imagery (importance:45 / dev:45)
- EMGBlend: Heterogeneity-Aware Self-Supervised Pretraining for Gesture and Force Decoding (importance:50 / dev:55)
- Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus (importance:55 / dev:45)
- What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation (importance:65 / dev:80)
- An Exploratory Replica-Overlap Probe of the Grokking Transition (importance:55 / dev:70)
- When Quantum Meets AI: Quantum Methods for Machine Learning and Machine Learning Methods for Quantum Systems (importance:55 / dev:65)
- From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs (importance:75 / dev:90)
- Teaching Reinforcement Learning and Humanoid Robotics to High-School Students: An Expert-Validated Curriculum Design on a Low-Cost Open Platform (importance:40 / dev:55)
- Interpretable AI plus Handheld, Portable Retinal Photographs: A Low-Cost Glaucoma Screening Solution for West Africa (importance:60 / dev:55)
- Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages (importance:60 / dev:75)
- Self-Supervised Combinatorial Optimization with Constraints via Frank-Wolfe (importance:60 / dev:80)
- Beyond Class Marginals: Bounding Rehearsal Gaps without Freezing Class Co-occurrence (importance:55 / dev:75)
- Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation (importance:55 / dev:55)
- Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models (importance:65 / dev:80)
- Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models (importance:70 / dev:80)
- You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs (importance:75 / dev:90)
- MorphoSHAP: Rethinking the Unit of Attribution in Explanation for Deep Visual Models (importance:65 / dev:80)
- CogenPVG: Cognitive-Enhanced Reflective Multi-Agent Framework for Persuasive Video Generation (importance:70 / dev:85)
- In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization (importance:60 / dev:80)
- Evaluating the Effectiveness of SechKAN on 1D Data (importance:50 / dev:70)
- Risk-Aware Online Conformal State Probing (importance:60 / dev:80)
- BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception (importance:70 / dev:85)
- Toward Responsible AI-Augmented Cyber Defense: Pattern Recognition, Defense-in-Depth, and the Case for Human-AI Collaboration (importance:75 / dev:85)
- Destination Support Restoration for Finite-Set Multimodal Trajectory Prediction (importance:60 / dev:80)
- Interweaving Marginals into Multivariate Sample Paths: Training-Free Dependence Construction for Probabilistic Time Series Foundation Models (importance:65 / dev:80)
- SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schr\"odinger Bridges (importance:65 / dev:80)
- Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models (importance:70 / dev:85)
- Compiling Sufficient Governance Context from Declared Losses and Reachable States: Exact Observation-Contract Synthesis with Cardinality and Cost Objectives (importance:70 / dev:90)
- Reciprocal Collaboration: how lessons from convergence in GLAMs can enhance interdisciplinary AI research (importance:50 / dev:45)
- REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States (importance:70 / dev:85)
- xWhyL: Causal Interactive Learning (importance:45 / dev:35)
- EMERGE: Resolution-Agnostic Point Cloud Generation with Equivariant Graph-Based Diffusion (importance:50 / dev:40)
- CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching (importance:30 / dev:35)
- Observing the Conduct of Systematic Reviews with Generative AI Support: An Experience Report from a Graduate Software Engineering Course (importance:45 / dev:45)
- Policy-Backed Selective Regeneration under Tainted Inter-Agent Communication (importance:55 / dev:65)
- TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models (importance:65 / dev:75)
- Certified Mechanistic Interpretability: Lifting Single-Input Findings to Bounded Neighbourhoods (importance:60 / dev:70)
- StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots (importance:65 / dev:70)
- TTTIR: Unlocking Instance-Specific State Evolution via Test-Time Training for Image Restoration (importance:45 / dev:50)
- MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning (importance:40 / dev:40)
- The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain (importance:60 / dev:65)
- Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models (importance:65 / dev:70)
- Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression (importance:55 / dev:70)
- Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems (importance:65 / dev:75)
- Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs (importance:60 / dev:70)
- WatchPoint: Executable User Feedback for Real-World Agentic Web Development (importance:70 / dev:80)
- When Unpaired Sets Support Shared-Corruption Calibration: Moment Geometry and Two-Sample Precision (importance:35 / dev:40)
- Reducing Hallucinations in Large Language Models Through Integrated Self-Verification and Retrieval-Augmented Generation (importance:65 / dev:75)
- ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap (importance:50 / dev:60)
- AIGC Video Detection based on the fusion of spatial-frequency-optical flow multimodal features (importance:60 / dev:70)
- On the security and privacy of LLMs in Mobility (importance:60 / dev:75)
- CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference (importance:70 / dev:80)
- TriWorldBench: A Tri-View Consistency Perspective on Embodied World Models (importance:55 / dev:65)
- Geometry-Aware Hyperbolic Residual Quantization (importance:50 / dev:65)
- TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling (importance:55 / dev:70)
- PACT: From Credit Assignment to Critic Alignment (importance:60 / dev:75)
- GitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement (importance:55 / dev:60)
- FairMean: Promoting Fairness in Distributed Learning under Label Poisoning Attacks (importance:50 / dev:65)
- MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation (importance:50 / dev:60)
- TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series (importance:55 / dev:70)
- QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation (importance:65 / dev:80)
- DeepFEAv2: Deep Learning for Transient Finite Element Analysis Beyond Structured Meshes (importance:50 / dev:60)
- Complementary Roles of Radiomics and Foundation Representations in Renal Cell Carcinoma Classification: A Comparative Study of 2D and 3D CT Encodings (importance:40 / dev:40)
- PP-Net: A Hybrid Physical-Prior Neural Network for Scattered Light Removal in Biomedical Images on Embedded Devices (importance:45 / dev:40)
- FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation (importance:70 / dev:85)
- Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing (importance:50 / dev:50)
- When Recursive Models Finish Computing (importance:55 / dev:75)
- Radiomics-Conditioned Modulation of RenalCLIP Features for Clear Cell Renal Cell Carcinoma Classification (importance:40 / dev:35)
- The Ethics of Artificial Intelligence in Military Operations (importance:60 / dev:40)
- Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model (importance:45 / dev:40)
- A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation (importance:55 / dev:70)
- Topology-Stratified Materials Discovery with A Flow-Based Generative Model (importance:50 / dev:50)
- The Disciplinary Language Transfer Problem: How Psychological Vocabulary Produces Governance Failures in AI Agent Deployment (importance:55 / dev:40)
- Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models (importance:55 / dev:65)
- Towards Hierarchical GNNs for multi-grid power flow: generalization across operating scenarios (importance:45 / dev:55)
- Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference (importance:60 / dev:80)
- Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models (importance:65 / dev:80)
- A Spectral Theory of Grokking: Weight Decay induces Feature Learning (importance:60 / dev:75)
- From Alignment to Access Control: A Framework for GenAI Policy Enforcement (importance:65 / dev:75)
- Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation (importance:65 / dev:85)
- Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning (importance:65 / dev:80)
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning (importance:65 / dev:80)
- TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits (importance:65 / dev:85)
- The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence (importance:65 / dev:80)
- Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows (importance:60 / dev:40)
- Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen (importance:70 / dev:85)
- FleXray: Universal Clinical X-ray Segmentation (importance:45 / dev:40)
- A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem (importance:75 / dev:85)
- SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue (importance:60 / dev:75)
- Stable Marriage Problems with Ties and Incomplete Preferences: An Empirical Comparison of ASP, SAT, ILP, CP, and Local Search Methods (importance:40 / dev:55)
- Decidable Reasoning About Time in Finite-Domain Situation Calculus Theories (importance:45 / dev:55)
- Small Language Models are the Future of Agentic AI (importance:75 / dev:80)
- Glucose-ML: A collection of longitudinal diabetes datasets for development of robust AI solutions (importance:50 / dev:45)
- EndoCogniAgent: Closed-Loop Agentic Reasoning with Self-Consistency Validation for Endoscopic Diagnosis (importance:50 / dev:45)
- Ultra Strong Machine Learning: LLM-Generated Explanations Do Not Yet Suffice for Teaching Humans Active Learning Strategy (importance:55 / dev:60)
- Navigating Taxonomic Expansions of Entity Sets Driven by Knowledge Bases (importance:45 / dev:55)
- Distributed Legal Infrastructure for a Trustworthy Agentic Web (importance:70 / dev:60)
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents (importance:75 / dev:85)
- SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory (importance:70 / dev:85)
- NIMO Controller: a self-driving laboratory orchestrator based on the Model Context Protocol (importance:75 / dev:85)
- Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads (importance:70 / dev:85)
- Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems (importance:70 / dev:85)
- Confidence Composition for Multiagent Language Model Systems (importance:70 / dev:85)
- Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation (importance:65 / dev:80)
- Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training (importance:50 / dev:55)
- SPINE: Bridging the Cyber-Physical Gap with Agentic AI (importance:70 / dev:85)
- Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models (importance:75 / dev:85)
- VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification (importance:60 / dev:75)
- DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making (importance:70 / dev:85)
- TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents (importance:70 / dev:85)
- Where Does Neural Advantage Arise in Continuous-Time Dynamic Graph Prediction? (importance:50 / dev:65)
- PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks (importance:50 / dev:70)
- Improving Constraint Models with LLM Agents (importance:70 / dev:85)
- BixBench3: Benchmarking AI agents on research-study-scale computational biology tasks (importance:60 / dev:70)
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance? (importance:60 / dev:80)
- MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games (importance:55 / dev:75)
- What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality (importance:65 / dev:80)
- Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation (importance:65 / dev:80)
- Potential of Artificial Intelligence Algorithms for Identification of Relevant Diagnostic and Prognostic Biomarkers of Early-Stage Liver Cancer (importance:50 / dev:40)
- ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training (importance:65 / dev:80)
- ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents (importance:70 / dev:85)
- Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure (importance:50 / dev:60)
- Pinocchio: Fast Uncertainty Estimates for Black-Box Language Models (importance:65 / dev:80)
- DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration (importance:60 / dev:75)
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control (importance:50 / dev:60)
- ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy (importance:55 / dev:70)
- How Can Incentives and Cut Layer Selection Influence Data Contribution in Split Federated Learning? (importance:50 / dev:65)
- FedNIA: Noise-Induced Activation Analysis for Mitigating Data Poisoning in Federated Learning (importance:60 / dev:75)
- BigO(Bench): Can LLMs Generate Code with Controlled Time and Space Complexity? (importance:70 / dev:85)
- Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors (importance:65 / dev:80)
- TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action (importance:48 / dev:65)
- OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots (importance:45 / dev:65)
- WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks (importance:70 / dev:85)
- Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint (importance:55 / dev:65)
- Real-time autonomous magnetic microrobot navigation across dynamic and biologically relevant environments (importance:35 / dev:55)
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique (importance:55 / dev:75)
- Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits (importance:38 / dev:65)
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy (importance:45 / dev:65)
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference (importance:65 / dev:75)
- Metamodel-Guided Model Generation with Layered Constraints (importance:55 / dev:85)
- STAR-VAE: A Scalable Latent-Variable Transformer for Controllable Molecular Generation (importance:45 / dev:75)
- Finding Kissing Numbers with Game-theoretic Reinforcement Learning (importance:45 / dev:65)
- Towards Synergistic Teacher-AI Interactions with Generative Artificial Intelligence (importance:55 / dev:45)
- Radiance-Field Guided Pretraining: Scaling Localization Models with Unlabeled Wireless Signals (importance:55 / dev:75)
- A Multimodal Large Language Model-Driven Framework for Context-Aware UAV Emergency Landing Site Selection (importance:55 / dev:75)
- SWE-Universe: Scale Real-World Verifiable Environments to Millions (importance:75 / dev:95)
- Discovering Data Manifold Geometry through Geometric Properties (importance:48 / dev:75)
- FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research (importance:55 / dev:75)
- Real Money, Fake Models: Deceptive Model Claims in Shadow APIs (importance:65 / dev:85)
- Spectral Overfitting in Noisy Linear Probing of Pretrained Representations (importance:48 / dev:75)
- Towards Effective Orchestration of AI x DB Workloads (importance:65 / dev:95)
- Separators in Enhancing Autoregressive Pretraining for Vision Mamba (importance:55 / dev:85)
- Seeing the imagined: latent functional alignment in visual imagery decoding from fMRI data (importance:45 / dev:65)
- A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Across the Memory Lifecycle (importance:75 / dev:95)
- Faithful Autoformalization via Roundtrip Verification and Repair (importance:55 / dev:85)
- Learning Dynamic Evidence Routes for Vision Transformer Probing (importance:48 / dev:75)
- LLM Ghostbusters: Surgical Package Hallucination Suppression via Adaptive Unlearning (importance:75 / dev:95)
- eXplaining to Learn (eX2L): Regularization Using Contrastive Visual Explanation Pairs for Distribution Shifts (importance:48 / dev:75)
- Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI (importance:65 / dev:85)
- The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions (importance:65 / dev:85)
- Unbiased Gradients, Moving Stability Boundaries: Exact Mini-Batch Geometry in Linear Self-Attention (importance:48 / dev:75)
- DDGAD: Disagreement-Driven Graph Anomaly Detection via Adapt-Then-Combine (importance:48 / dev:75)
- AgenticDiffusion: Multi-View Reasoning with View-Conditioned Diffusion Planning for Vision-Based UAV Navigation (importance:55 / dev:85)
- Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport (importance:45 / dev:65)
- When Good Verifiers Go Bad: Silent Negative Transfer in Verifier-Guided VLM Training (importance:55 / dev:85)
- FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming (importance:65 / dev:85)
- Explanation-Guided Medical Named Entity Recognition with Stability and Boundary Awareness for Atopic Dermatitis (importance:45 / dev:75)
- SOLAR: AI-Powered Speed-of-Light Performance Analysis (importance:65 / dev:95)
- Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits (importance:55 / dev:85)
- ReasonLab: A Controlled and Auditable Evaluation of Prompting Techniques for Multiple-Choice QA (importance:55 / dev:85)
- From Plausible to Actionable: A Position on LLM Self-Explanations (importance:55 / dev:85)
- A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance (importance:60 / dev:75)
- Semi-Automated Detection of Gaps in LLM Security Knowledge (importance:65 / dev:85)
- Parameter-Free Dynamic Regret under Heavy-Tailed Noise (importance:35 / dev:65)
- CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation (importance:55 / dev:65)
- HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation (importance:65 / dev:95)
- Are Concept Bottleneck Models Effective as Decision-Support Systems? (importance:55 / dev:75)
- Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation (importance:55 / dev:85)
- GVS5H: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance (importance:75 / dev:95)
- Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict (importance:55 / dev:75)
- Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention (importance:48 / dev:75)
- HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference (importance:75 / dev:95)
- Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory (importance:55 / dev:85)
- When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents (importance:65 / dev:85)
- Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings (importance:45 / dev:75)
- VERPO: Verified Evidence Regularized Policy Optimization (importance:65 / dev:85)
- Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint) (importance:48 / dev:75)
- Rice's Theorem under Self-Modification: Elevation Operators and a Normal Form (importance:45 / dev:75)
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement (importance:65 / dev:75)
- Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning (importance:55 / dev:85)
- SkillAtlas: An Attack Trace Library for Agent Skills (importance:75 / dev:95)
- Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection (importance:55 / dev:85)
- ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation (importance:48 / dev:45)
- Planning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs (importance:55 / dev:85)
- Models as Governed Interfaces for AI-Native MBSE: Read-Side Adequacy and Write-Side Admissibility (importance:55 / dev:85)
- PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control (importance:55 / dev:75)
- Efficient Nash Equilibrium Computation for Cybersecurity Games (importance:50 / dev:75)
- A Functional Pilot for Certified Freshness-Aware Semantic--Spatial Range Retrieval (importance:48 / dev:75)
- Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs (importance:55 / dev:85)
- Quantifying Overclaiming Propensity in Frontier LLM Agents (importance:75 / dev:95)
- Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents (importance:55 / dev:85)
- Beyond Task Completion: Training Capable and Safe Computer-Use Agents (importance:75 / dev:95)
- On The Robustness-Resolution Tradeoff In Temporal Quantization Of Event Streams (importance:48 / dev:75)
- AffordanceWAM: Affordance-Aware Joint World-Action Modeling for Robot Manipulation (importance:55 / dev:85)
- Parameterized Dense-Sparse Fusion for Hybrid Retrieval: Tuning a Rank-Score Mix on BEIR SciFact with Qdrant (importance:55 / dev:85)
- Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs (importance:55 / dev:85)
- RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents (importance:75 / dev:95)
- MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents (importance:75 / dev:95)
- Estimating Accurate Hand Pose in Camera Space with Vision Transformer (importance:48 / dev:75)
- ActGov: Governing LLM Agent Actions via Policy-Constrained Validation (importance:75 / dev:95)
- VPRune: Efficient Training-free Pre-LLM Visual Token Pruning (importance:65 / dev:85)
- Touch2Robot: Robot Touch in the Human Demonstration Loop (importance:55 / dev:75)
- Mobile Imaging Solutions for Medical Diagnosis: Trends and Applications (importance:45 / dev:55)
- Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI (importance:75 / dev:95)
- DolphinBench: Mapping the Pareto Frontier of Agent Memory (importance:75 / dev:95)
- Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow (importance:55 / dev:85)
- Dual-GNN Multilevel Coarsening for Maximum Independent Set (importance:45 / dev:75)
- Learning Neural Feedback Linearization for Data-driven Systems via Augmented Lagrangian (importance:45 / dev:75)
- Correcting Within-Group Self-Selection Bias in Prioritized Replay (importance:48 / dev:75)
- Topological Signal Processing With Unoriented Operators (importance:35 / dev:65)
- Spatiotemporal Kronecker Covariance Neural Networks (importance:48 / dev:75)
- MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data (importance:45 / dev:65)
- Concept Drift from a Causal Perspective (importance:55 / dev:75)
- Extending FunctionGemma for Practical On-Device Mobile Function Calling (importance:65 / dev:85)
- PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics (importance:45 / dev:75)
- Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models (importance:55 / dev:85)
- Learning Defensive Policies against Diverse Inference Attacks for Smart Meter Privacy (importance:55 / dev:75)
- Continuous Optimization for p-adic Models (importance:35 / dev:65)
- SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling (importance:45 / dev:65)
- Rewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs (importance:55 / dev:85)
- Efficient Cost-Aware LLM Evaluation via Bayesian Bandit Gittins Indices (importance:30 / dev:60)
- Targeted Review for AI-Assisted Biodiversity Surveys: Active Continuous-Score Occupancy Modeling (importance:25 / dev:30)
- When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds (importance:25 / dev:50)
- Marginal Log-Likelihood Increments under Dirichlet-Smoothed Markov Estimation (importance:15 / dev:40)
- Graph Domain Adaptation Does Not End with Representation Learning (importance:35 / dev:55)
- Fully Byzantine-Resilient Multi-Agent Reinforcement Learning (importance:40 / dev:65)
- Signed Graph Pre-Training and Prompt Learning (importance:30 / dev:50)
- Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations (importance:35 / dev:55)
- Minimal Recurrent Behavioral Memory for Imitation under Partial Observability (importance:25 / dev:45)
- Disentangling Heterogeneous Traffic Dynamics for Multi-Step Traffic Forecasting via Adaptive Spectral Decomposition (importance:20 / dev:45)
- A Lightweight Plastic-Memory Framework for Graph Few-Shot Class-Incremental Learning (importance:30 / dev:50)
- Latest Exact Match Attention (importance:40 / dev:65)
- Auditing Proxy-Based Validation Across Text Spans (importance:30 / dev:50)
- Multi-View Fair Clustering Guided by Cross-View Sensitive Information Discrepancy (importance:30 / dev:50)
- CacheDyG: Decoupling Temporal Propagation for Efficient Dynamic Graph Learning (importance:35 / dev:60)
- Protocol before progress: leakage-aware evaluation of AIS trajectory prediction (importance:25 / dev:45)
- Gaussian Flow-Matching Schedules: Implications for Sampling and Training (importance:35 / dev:55)
- Neural Approximation by Function Composition: Rigidity and Doubly Exponential Convergence (importance:30 / dev:50)
- AURA: Angular Update Rate Adaptation for training complex-valued neural networks (importance:30 / dev:55)
- Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement (importance:45 / dev:65)
- Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL (importance:30 / dev:55)
- Exploring Solver-Level Warmstarting for Neural Network Verification (importance:35 / dev:60)
- GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression (importance:40 / dev:65)
- Theory for groupoid equivariant neural networks: an approach for steerable CNNs on bounded domains (importance:30 / dev:55)
- The Dynamics of Quasiregular Neural Learning (importance:30 / dev:50)
- BOBA: Dynamic Bayesian Optimization through Bayesian Active Inference (importance:35 / dev:60)
- MICRO: Multi-Fidelity Active Search for Severe Error Discovery (importance:30 / dev:55)
- Towards Adaptive Federated Graph Clustering: A Global Community-aware Contrastive Learning-based Approach (importance:30 / dev:55)
- FuncCode: Compressing Kolmogorov--Arnold Networks in Function Space with Hardware-Aware Quantization (importance:35 / dev:60)
- Fast Matrix Multiplication in fp8: Certified Coefficient Optimization and Measured Error (importance:35 / dev:65)
- Margin-Drop Coordinates for Cross-Budget Robustness Evaluation (importance:30 / dev:55)
- CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning (importance:35 / dev:55)
- From Risk Scoring to Risk Allocation: A Density-Driven Framework for Diverse Monitoring in Multi-Agent Systems (importance:35 / dev:60)
- Block-Level Weight-Space Structure Persists Under Post-Training: An Empirical Study Across LLM Families (importance:40 / dev:65)
- Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery (importance:40 / dev:65)
- Activation-Energy Pruning for Spiking Neural Networks: Unsupervised Personalization via Spike-Count Saliency (importance:30 / dev:55)
- Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity (importance:35 / dev:60)
- Partially Observed Sparse Graphs: The Unknown Sampling Rate is a Tail Index (importance:25 / dev:50)
- Bridging the Data Gap: Digital Twin as a New Paradigm for AI-based Radio Sensing (importance:35 / dev:60)
- Beyond Imitation: Auditing the Recoverability of Reasoning in Distilled Models (importance:40 / dev:60)
- MSA-CITE: A Co-Adapted LoRA Specialist Ecology for Fixed-Budget Small-Model Inference (importance:40 / dev:65)
- High-Order Liquid Evidence Modeling for Continuous and Subtle GNSS Spoofing Detection in Autonomous Driving (importance:30 / dev:55)
- Can You Delete a Year of Market Data? Machine Unlearning Against Exact Retraining Oracles (importance:40 / dev:60)
- Information-Theoretic Decoupled Prompt Tuning for Continual Learning (importance:35 / dev:55)
- Mode Collapse Is Cheap to Detect: A Ground-Truth-Free Pre-Flight Check for Neural Samplers (importance:30 / dev:55)
- JAMPR+/L2D: scalable neural heuristic for constrained vehicle routing problems in dynamic environment (importance:30 / dev:55)
- On the Effect of Bit-Level Parameter Perturbations in Machine Learning and Deep Learning Models (importance:30 / dev:55)
- Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting (importance:25 / dev:45)
- PreGS: A Parameter-Transfer-Based Multi-Expert Graph Neural Network for Node Classification (importance:30 / dev:55)
- Disaggregated Quantization: Specializing LLM Prefill and Decode (importance:45 / dev:70)
- Learning to Defer with Guidance on Real World Medical Data (importance:30 / dev:50)
- Double Descent and Malign Overfitting in Diffusion Models (importance:40 / dev:60)
- OMatG-flash: An All-Atom Flow Map with Reinforce Adjoint Matching for Scalable Materials Discovery (importance:30 / dev:55)
- One-Step Generative Surrogate Models via Block-Triangular Joint Drifting (importance:35 / dev:60)
- Can We Predict Anomaly Detection Performance from Embedding-Space Geometry? (importance:30 / dev:55)
- Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy (importance:35 / dev:60)
- Notes on Fourier-Bessel wavelets (importance:15 / dev:40)
- Label-Efficient Learning for Ground-Based Sky-Image Classification: A Benchmark of Transfer Learning, Active Learning, and Pseudo-Labeling on GCD (importance:25 / dev:50)
- MAGIC: Mixed-Granularity Agent Graphs via Incremental Construction with Dense-Reward Reinforcement Learning (importance:40 / dev:65)
- EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations (importance:35 / dev:65)
- From IceCube to IT-Sphere: A Hybrid Quantum-Classical GNN for Banking IT Root Cause Analysis (importance:30 / dev:60)
- Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment (importance:40 / dev:55)
- When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits (importance:30 / dev:55)
- What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus (importance:35 / dev:55)
- Same Quantity, Different Answer: Numerical Representation Invariance in Language Models (importance:35 / dev:55)
- From Tone to Trajectory: Continuous Sentiment and the Shape of Monetary Policy Communication (importance:25 / dev:40)
- What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content (importance:40 / dev:60)
- End-to-End Quantum Semantic Communication with Variational Quantum Neural Networks (importance:30 / dev:60)
- SPARC: SuperPixel-Aware Region Contrastive Learning for Self-Supervised Dense Prediction (importance:30 / dev:55)
- FREESIA: Covariance-Aware Posterior Transport for Expressive and Scalable Data Assimilation (importance:30 / dev:60)
- Calibration Count Reuse: Validity Does Not Determine Efficiency (importance:25 / dev:50)
- The Informational Content in Lepto-Variance and Its Relation to Higher Moments (importance:20 / dev:45)
- Variational objectives for amortized Bayesian inference in inverse problems: The role of posterior conditioning (importance:30 / dev:60)
- Empirical Auditing of Edge-Private Graph Generators (importance:35 / dev:60)
- GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry (importance:30 / dev:60)
- Learning from Humans for Proactive Assistance in Human-Robot Collaborative Transport (importance:30 / dev:60)
- SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation (importance:45 / dev:65)
- Penalized Nonreversible Langevin for Constrained Sampling (importance:25 / dev:55)
- Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics (importance:20 / dev:50)
- Learnable Classifier-Free Guidance Null Embeddings for Enhanced Controllable Speech Synthesis (importance:30 / dev:60)
- WeightBridge: An Efficient Weight Transfer Library for Reinforcement Learning (importance:40 / dev:70)
- MIND the Gap: A Geographic Implicit Neural Representation with Adjustable Spatial Scale (importance:30 / dev:60)
- SAM-V: Geometry-Aware Segment Anything for Multi-View Instance Segmentation (importance:35 / dev:60)
- FAST-ML: A Hybrid Physics-Machine Learning Framework for Tropical Cyclone Intensity Forecasting (importance:30 / dev:60)
- Matryoshka attribution: Learning to attribute language model outputs to representations and weights (importance:40 / dev:65)
- Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning (importance:30 / dev:60)
- HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery (importance:35 / dev:65)
- Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis (importance:25 / dev:55)
- Generalized Deep Regression for Repeated Measurements (importance:25 / dev:55)
- Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures (importance:45 / dev:75)
- CODA: Depth-Aligned Scene Completion and Object Decomposition from a Single RGB-D Image (importance:30 / dev:60)
- On the Gradient Heterogeneity Dynamics of Adversarially Robust Federated Regression (importance:30 / dev:60)
- Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression (importance:30 / dev:60)
- Graded Representation Theory of Equivariant Neural Networks (importance:30 / dev:60)
- Statistical Gains from Looped Estimation under Parameter Budgets (importance:25 / dev:55)
- Beyond Reconstruction Error: Analytical and Data-Driven Action Tokenization for Autoregressive Vision-Language-Action Models (importance:35 / dev:65)
- Visual Jev: Accurate and Efficient Decisions from Shared Visual Context (importance:30 / dev:60)
- Conditional Tensor Diffusion: Distributional Counterfactual Learning and Inference (importance:30 / dev:60)
- Bridge of $\Psi$'s: Quantum Circuit Optimization with Schr\"odinger Bridges (importance:30 / dev:65)
- Faithful Faithfulness Evaluations: Challenges & Pitfalls Learned from a Breast MRI Case Study (importance:30 / dev:55)
- Hyperbolic Restricted Boltzmann Machine Neural Quantum State (importance:5 / dev:15)
- Optimizing Denoising Trajectories in dLLMs: A Lightweight Evolutionary Heuristic Approach (importance:35 / dev:40)
- Differentiable Policy Transport over Multi-Layer Network Feasibility Geometry (importance:10 / dev:35)
- One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data (importance:20 / dev:25)
- TailSpec-EASE: Knowledge-Graph-Regularized Linear Recommendation for Web Long-Tail Discovery (importance:25 / dev:30)
- PatchKV: Efficient KV Cache Recovery for Dynamically Edited LLM Contexts (importance:50 / dev:65)
- Damage Predicts Recovery: When Calibration Data Matters in Compressing Financial LLMs (importance:30 / dev:45)
- Sample-Smooth Spaces: A Convenient Category for Differentiable Probabilistic Programming (importance:5 / dev:20)
- Learning to Fluctuate: Statistical Foundations for Causal Tabular Pretraining (importance:25 / dev:35)
- Target alignment, dilution and forecast selection when cross-sectional forecasts share a common target (importance:5 / dev:10)
- Error Bounds for Statistical Estimators in BTL Model with Parametric Multivariate Utility Functions (importance:5 / dev:5)
- HYDRA: Proactive Android Malware Drift Adaptation via Hierarchical Graph Contrastive Learning (importance:40 / dev:50)
- On the Lexical Superstition of Large Language Models for Code Comprehension: Re-evaluation on Code of Low Lexical Quality (importance:50 / dev:75)
- SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA (importance:15 / dev:25)
- A Practical Guide on Graphical Model Validation (importance:15 / dev:20)
- Deep Generative Crystal Structure Prediction: A Benchmark Study and a Controlled Test of Prototype Dependence (importance:25 / dev:30)
- Polyak-Type Extragradient Methods for Monotone Root-Finding Problems (importance:10 / dev:15)
- GTR: Gated Token Recurrence for Efficient Dense Prediction (importance:40 / dev:55)
- Unlocking Cross-Scenario Physical Layer Security: A Mixture-of-Experts Framework with Generative Diffusion Models (importance:20 / dev:30)
- Foundation model embeddings capture pre-diagnostic changes on screening mammograms (importance:35 / dev:25)
- MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction (importance:30 / dev:20)
- On Basis Function Selection for Sparse Gaussian Process Regression (importance:15 / dev:25)
- Statistical Rates for Entropic Optimal Transport in the Discrete to SubGaussian Regime (importance:5 / dev:10)
- Discovery-Driven Integration of Disjoint Tables via Text (importance:35 / dev:50)
- PROSWIN: Probabilistic Solar Wind Speed Forecasting Using Deep Distributional Regression From Solar Images (importance:20 / dev:25)
- When are bosonic Gaussian states classical to learn? (importance:5 / dev:10)
- Optimal Sequential Annotations for Off-Policy Evaluation (importance:30 / dev:40)
- Diffusion-Induced Spatial Attention Overlapping Community Detection (importance:20 / dev:35)
- Automatic depth-based local center clustering via $\beta$-integrated local depth and adaptive grouping (importance:15 / dev:30)
- A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing (importance:10 / dev:25)
- DeepSPoC: A Deep Learning Based Sequential Propagation of Chaos (importance:15 / dev:30)
- Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies (importance:20 / dev:20)
- Optimizing Canaries for Privacy Auditing with Metagradient Descent (importance:50 / dev:65)
- Transport-Coupled Bayesian Flows for Molecular Graph Generation (importance:30 / dev:30)
- Robust Photoplethysmography Signal Denoising via Mamba Networks (importance:20 / dev:25)
- Simulation-free Structure Learning for Stochastic Population Dynamics (importance:15 / dev:30)
- CID: Measuring Feature Importance Through Counterfactual Distributions (importance:40 / dev:50)
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning (importance:60 / dev:75)
- Linear probing enables Ship-Radiated Noise recognition with pretrained audio embeddings (importance:15 / dev:20)
- Relative Wasserstein Angle and the Problem of the $W_2$-Nearest Gaussian Distribution (importance:5 / dev:10)
- Quantum Model Parallelism for MRI-Based Classification of Alzheimer's Disease Stages (importance:20 / dev:25)
- Communication-Efficient Byzantine-Robust Federated Conformal Prediction via Partial Sharing (importance:45 / dev:60)
- RepUCB: Representation Learning-Based UCB for Heterogeneous Multi-Task Linear Bandits (importance:20 / dev:35)
- SPLICE: Latent Diffusion over JEPA Embeddings for Conformal Time-Series Inpainting (importance:35 / dev:50)
- Event-Based Early Warning of Vineyard Disease Risk from Environmental Time Series (importance:15 / dev:20)
- Practical Scaling Laws: Converting Compute into Performance in a Data-Constrained World (importance:70 / dev:75)
- Tight Sample Complexity Bounds for Entropic Best Policy Identification (importance:15 / dev:30)
- CAffNet: Hard Constraint-Affine Neural Networks (importance:35 / dev:50)
- MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training (importance:55 / dev:70)
- Refit the Probe: Single-Direction Ablation Is Not a Necessity Test (importance:35 / dev:55)
- Sharp First-Order Lower Bounds for Higher-Order Smooth Nonconvex Optimization (importance:10 / dev:20)
- Evidential Fusion Network for Multimodal Survival Prediction under Missing Modalities (importance:20 / dev:20)
- Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization (importance:35 / dev:50)
- Converge to Surprise: Evolutionary Self-supervised Image Clustering (importance:30 / dev:45)
- Low-Rank Attention Residuals (importance:50 / dev:70)
- Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments (importance:20 / dev:35)
- Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion (importance:20 / dev:20)
- Sharp Characterization of Bias in Post-Bandit Inference (importance:20 / dev:30)
- Task- and dataset-specific information in protein language models (importance:25 / dev:30)
- Orthogonal JEPA: Factorized Predictive States for Latent World Models (importance:50 / dev:65)
- Risk-Conditioned Fine-Tuning of Large Language Models (importance:55 / dev:70)
- Exact-Form Regret for Gradient Descent, Mirror Descent and Follow-the-Regularized-Leader (importance:10 / dev:20)
- LiveProBench: Can Streaming Video Models Really Interact Like Humans? (importance:40 / dev:55)
- GRPO-QPS: Target-Preserving Reinforcement Learning for Quantum Posterior Sampling (importance:15 / dev:25)
- Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation (importance:45 / dev:65)
- Continuous Delayed-Memory Stochastic Gradient Descent and Continuous-Time Reinforcement Learning from History of Astrophysical Time Series Studies (importance:15 / dev:25)
- Ranking Competing geologic interpretations via foundation-model-assisted generative hydrologic inversion (importance:15 / dev:25)
- Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data (importance:30 / dev:40)
- $\lambda$-Controlled GRPO: Turning Flow-Matching Ratio Instability into a Budgeted Resource (importance:40 / dev:60)
- PAGE: Partition-Aware Gated KV-Cache Eviction (importance:55 / dev:75)
- StepKV: Step-Aware KV Cache Compression for LLM Agents (importance:60 / dev:80)
- Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation (importance:35 / dev:50)
- Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting (importance:35 / dev:50)
- A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting (importance:45 / dev:60)
- Lifelong Learning of Video Diffusion Models From a Single Video Stream (importance:45 / dev:65)
- Polynomial Scaling is Possible For Neural Operator Approximations of Structured Families of BSDEs (importance:15 / dev:30)
- The Challenge of Identifying the Origin of Black-Box Large Language Models (importance:50 / dev:65)
- Riemannian Optimization on Tree Tensor Networks with Application in Machine Learning (importance:10 / dev:25)
- Finite Topological Space Filtrations: A Topological Framework for Data Analysis (importance:15 / dev:25)
- Exact and Approximate Range Queries in Ball Mapper (importance:15 / dev:25)
- Learning to bin: differentiable and Bayesian optimization for multi-dimensional discriminants in high-energy physics (importance:15 / dev:25)
- Text-only adaptation in LLM-based ASR through text denoising (importance:45 / dev:60)
- VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean (importance:50 / dev:75)
- Conditional Distributional Treatment Effects: Doubly Robust Estimation and Testing (importance:15 / dev:25)
- Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction (importance:30 / dev:35)
- Unified Multimodal Uncertain Inference (importance:40 / dev:55)
- The Virtue of Sparsity in Complexity (importance:15 / dev:20)
- Flow Matching for Count Data (importance:35 / dev:50)
- Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising (importance:20 / dev:35)
- Financially Guided Deep Portfolio Optimization (importance:25 / dev:40)
- Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis (importance:25 / dev:30)
- Rethinking Post-Hoc Calibration in Semantic Segmentation (importance:35 / dev:55)
- Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification under Foundation-Model Pretraining (importance:30 / dev:50)
- Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging (importance:30 / dev:50)
- Chaos Is a LADDER: Domain Generalization Beyond Invariance via Reweighting (importance:40 / dev:60)
- A Quantum/Classical Example Oracle Separation for Making Things Up (importance:10 / dev:15)
- Interpretable AI with Local Distillation (importance:45 / dev:65)
- Disciplined Bilevel Programming (importance:35 / dev:60)
- Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules (importance:25 / dev:35)
- Memoization Without Keys: Compact, Out-of-Core Tables for Functions of Sorted Arguments (importance:35 / dev:70)
- Density-Ratio Rescoring for Imbalanced Classification Using Raking Duals and Classifier Scores (importance:20 / dev:30)
- Poisson Exchange Beyond Submodularity: Effective Approximation Algorithms for Offline and Online Subset Selection over Matroids (importance:15 / dev:20)
- TRACTOR Benchmark for Evaluating C to Rust Translators (importance:50 / dev:70)
- TAILOR: Template-Preserving Augmentation for Long-Tailed Log Parsing (importance:20 / dev:40)
- The Vocabulary of Flaky Tests in Swift (importance:30 / dev:60)
- Dynamic Conformance Testing of WebGPU Through Specification-Driven Mutation (importance:25 / dev:50)
- Evaluating Shaker for Flaky Test Detection in Python Projects (importance:30 / dev:60)
- An Empirical Analysis of Cross-OS Portability Issues in Python Projects (importance:20 / dev:50)
- Understanding Maintenance and Support in a Community-Driven Scientific Workflow Ecosystem: A Cross-Space Study of Galaxy (importance:15 / dev:40)
- What Was Once Learned May Need to Be Unlearned: Machine Unlearning for Deprecated API Knowledge in Large Language Models (importance:60 / dev:70)
- Confidence-Guided Cross-Modal Knowledge Transfer for Multimodal Anomaly Detection in Microservice Systems (importance:25 / dev:50)
- When Should Dependency Updates Invoke Repair Agents? A Lightweight Routing Study (importance:50 / dev:80)
- On Behavioral Alignment of Model-Code and Human-Code Understandability via Behavioral Proxies (importance:40 / dev:70)
- Post-Hoc Attention Steering of Large Language Models for Robust Code Understanding under Obfuscation (importance:40 / dev:60)
- Why Do LLMs Fail at OCL Generation? A Graph Reasoning Perspective (importance:30 / dev:50)
- A Large-Scale Longitudinal Study of Multi-CI Service Adoption (importance:35 / dev:70)
- CANcept: Model-based CAN Traffic Generation and Manipulation (importance:20 / dev:40)
- From Approval to Execution: Reconstruction-Aware Repair Analysis for LLM-Agent Software (importance:60 / dev:75)
- OSFoundry: Building and Evolving Operating Systems with Specification-Guided Agents (importance:70 / dev:80)
- Modular Composition of Inductive Types Using Lean Meta-programming (importance:25 / dev:60)
- Testing and Learning Symbolic Finite State Machines (importance:20 / dev:40)
- Towards Systematic Qualification of Vision-Language Models for Automotive Perception Systems (importance:40 / dev:50)
- Towards Effective Black-Box Adversarial Attacks on Deep Code Models via Structural and Identifier Perturbations (importance:50 / dev:70)
- Design and Evaluation of a Controlled Post-Alert Incident Orchestration and Response Subsystem Using a Rule Engine and a Local Large Language Model (importance:40 / dev:65)
- LLM-Based Repair of Static Nullability Errors (importance:50 / dev:75)
- Petrify: Petri-net Based Analysis of Concurrency Properties in Java Bytecode (importance:30 / dev:60)
- XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery (importance:50 / dev:60)
- AutoSQL: Extracting SQL Templates from Imperative ORM Code in Large-Scale Repositories (importance:40 / dev:70)
- Isabelle/STARK: A Formalization of zk-STARK in Isabelle/HOL (importance:30 / dev:50)
- The GNOME LLM Policy That I Want (importance:10 / dev:30)
- I want my mesh networks to be signed, not encrypted (importance:5 / dev:20)
- Abandoning Scientific Linux Was a Mistake (importance:5 / dev:20)
- KDE for People (importance:3 / dev:15)
- No Sloptober (importance:8 / dev:30)
- The Zig Journey (importance:10 / dev:40)
- Plain-text files are at risk (importance:8 / dev:25)
- Cheaper LLM labelling (importance:30 / dev:50)
- That About Wraps It Up for Stock Mac UI (importance:5 / dev:15)
- Do not let your type system reason about aliasing in your programming language (importance:20 / dev:50)
- How to talk about "AI" without adding to the anthropomorphization (importance:15 / dev:20)
- Parsing JSON Objects without intermediate ASTs (importance:20 / dev:50)
- Why didn't anybody tell me about Redis hash slots? (importance:15 / dev:40)
- Latest BGP hijack targets hosting software vendor (importance:50 / dev:60)
- Sandboxing with minimal effort (importance:35 / dev:70)
- Jev-powered autocorrection (importance:40 / dev:50)
- Looking forward to Git 2.56 - and 3.0 (importance:25 / dev:70)
- Fearless SIMD v1.0 is here (importance:20 / dev:60)
- GitHub Actions leaking secrets when Miri output is cached (importance:60 / dev:75)
- Serious editors are a commitment (at least for me) (importance:10 / dev:30)
- Re: Suggestions on implementing an efficient instruction set simulator in LuaJIT2 (2011) (importance:5 / dev:30)
- The Philosopher Plush — a DIY AI toy that runs 100% local (importance:0 / dev:30) (いいね相当スコア: 0)
- Google Gemini Adds Webflow Support for AI Content and CMS Workflows (importance:0 / dev:40) (いいね相当スコア: 0)
- TypeSafe AI Jev explained: examples, use cases, and LLM comparisons (importance:0 / dev:70) (いいね相当スコア: 0)
- Road to State Machines Part I (importance:0 / dev:70) (いいね相当スコア: 0)
- Everything i built (and cut) shipping a multiplayer app generator in 4.5 days (importance:0 / dev:70) (いいね相当スコア: 0)
- # From Idea to Venture The Role of a Venture Studio (importance:0 / dev:20) (いいね相当スコア: 0)
- We Bought the Agents. Then the Blank Canvas Ate Our Credits. (importance:0 / dev:60) (いいね相当スコア: 0)
- Cross-Harness Tool Parity: Build One Custom MCP Tool, Deploy It Everywhere (importance:0 / dev:85) (いいね相当スコア: 0)
- Jev vs. LLMs: What Happens When You Strip Text Generation Out of a Language Model (importance:0 / dev:75) (いいね相当スコア: 0)
- Google Gemini Enterprise Adds Airtable Connector for Search and Workspace Actions (importance:0 / dev:35) (いいね相当スコア: 0)
- Fail the Merge When an Agent Deletes or Weakens Tests (importance:0 / dev:80) (いいね相当スコア: 0)
- Kavach: building a fraud investigator that knows when a risk score is lying (importance:0 / dev:60) (いいね相当スコア: 0)
- I Compared 5 LLM Gateway Tools for Real-World Production Use (importance:0 / dev:80) (いいね相当スコア: 9)
- Claude Code Pricing — What You're Actually Paying For ☕️ (importance:0 / dev:60) (いいね相当スコア: 0)
- Anatomy of a Voice Agent: VAD, STT, LLM, TTS and Why WebRTC Matters (importance:0 / dev:75) (いいね相当スコア: 0)
- What Grok catches, what Codex catches, and what the pair costs us (importance:0 / dev:75) (いいね相当スコア: 1)
- LLM Security Is Not Just Prompt Injection: Understanding the Full Attack Surface (importance:0 / dev:80) (いいね相当スコア: 0)
- Designing Tool-Scoped Subagents to Prevent Context Bloat in Agent Runtimes (importance:0 / dev:85) (いいね相当スコア: 0)
- Pin Fixture, Grader, and Slot Before You Diff (importance:0 / dev:70) (いいね相当スコア: 0)
- I graded my memory server next to Mem0's published answers. It's a tie — and here is how Mem0 gets to 92.5. (importance:0 / dev:70) (いいね相当スコア: 0)
- Your LLM eval set is quietly certifying your bugs (importance:0 / dev:75) (いいね相当スコア: 0)
- Temperature 0 Isn't Deterministic: Batch Invariance in LLM Inference (importance:0 / dev:80) (いいね相当スコア: 0)
- OPUS 5.5 IS THE NEW 4.6! (importance:35 / dev:50)
- Opus 5.5 vs GPT-6 Sol: 3D Pelican riding bike test in Blender (importance:25 / dev:40)
- AI still has its moments!! (importance:5 / dev:10)
- Claude is BACK! (importance:15 / dev:20)
- Anthropic chief scientist Jared Kaplan warned AI training could trigger an intelligence explosion (importance:60 / dev:30)
- Stupid question - if Opus 5.5 is better than Fable 5.1, why would we ever use Fable 5.1? (importance:20 / dev:40)
- Opus 5.5 creates a train journey drawn entirely in JavaScript (importance:10 / dev:20)
- I do scientific research and Opus5.5 refuses to touch anything I’ve been working on. (importance:20 / dev:20)
- Vibe coding Minecraft: January this year vs. today (importance:25 / dev:70)
- Opus 5.5 built me a website that turns any photo into one-line art. It also films the line being drawn. (importance:15 / dev:40)
- Opus 5.5: First impressions by a trained philosopher (importance:20 / dev:30)
- Created a Rube Goldberg machine in crayon with Opus 5.5 (importance:5 / dev:10)
- Opus 5.5 built me a website that turns any photo into ASCII art (importance:10 / dev:30)
- Opus 5.5 is 40% cheaper while being 30% faster than opus 5. (importance:50 / dev:70)
- Holy shit, it refuses to eat usage. (importance:20 / dev:40)
- Claude opus 5.5 WOW (importance:10 / dev:20)
- haven't raged once since opus 5.5 (importance:3 / dev:5)
- 2D Game Art Tutorial Video By Claude (Opus 5) (importance:10 / dev:30)
- Dear Anthropic please do not nerf (importance:5 / dev:10)
- Has anyone used their reset yet? (importance:5 / dev:10)
- Opus 5.5 in Claude Code is crazy fast, especially at spotting UI bugs (importance:25 / dev:70)
- Lighthouse SVG: Opus 5.5 vs Sol 6 vs Astra vs Fable 5.1 (importance:30 / dev:50)
- ChatGPT SOL 6 Used Local Qwen for heavy lifting! (importance:35 / dev:50)
- Finally moved off claude code to OSS models on opencode (importance:20 / dev:60)
- Anyone here want some free help getting started building something? (importance:5 / dev:20)
- Does building your own tools makes sense? (importance:15 / dev:50)
- How do you split AI models across ideation, math, and coding? (importance:25 / dev:70)
- Will your next app be an iMessage/WhatsApp bot? (importance:20 / dev:50)
- Should you build your own EU VAT number validation against VIES or outsource it? (importance:5 / dev:30)
- LLM prompt injection testing at work just nuked our client demo and I feel sick (importance:45 / dev:65)
- Coming from Claude Code — how do you actually work with Codex's 258k context window? (importance:40 / dev:70)
- How enable subagent mode in ChatGPT Pro6 again (importance:15 / dev:55)
- The agent exited cleanly with status 0, did nothing, and reported success (importance:50 / dev:75)
- Companies bragging about "3x productivity" from AI coding, anyone else hearing the other side of that story? (importance:50 / dev:65)
- Should an AI coding agent ever be allowed to merge its own PR? (importance:55 / dev:70)
- Need help bypassing CAPTCHA while using Claude (importance:20 / dev:60)
- I think AI coding made it too easy for me to keep changing my app (importance:35 / dev:60)
- Mods: can we do something about half the forum getting filled with these advertising posts for Jev? (importance:5 / dev:0)
- MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today. (importance:50 / dev:65)
- Jev isn't new tech. Its marketing targets people who think AI started with LLMs. (importance:35 / dev:40)
- Pirate Face - pirate bay for LLMs (importance:25 / dev:20)
- this is not even a competition at this point ... this is embarrassing (importance:3 / dev:0)
- Cost of intelligence is dropping fast (importance:45 / dev:35)
- MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam (importance:50 / dev:60)
- GGUFs in transformers natively! (importance:55 / dev:80)
- apple/LensVLM-9B · Hugging Face (importance:60 / dev:70)
- BFL releases FLUX 3 Action: a 7B robot model (importance:55 / dev:60)
- Streaming Nemotron 3 Diarization (importance:45 / dev:75)
- Engram gone wild! 2b model update... (importance:30 / dev:55)
- Perhaps the highest quality mainline quants of Qwen3.8 27B? (importance:45 / dev:75)
- Pi agent qwen 3.8 flash next plays Baldur's Gate 2 (importance:25 / dev:65)
- Most powerful harness for Qwen 3.8? (importance:35 / dev:75)
- GUI harness [video] (importance:30 / dev:70)
- DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes) (importance:45 / dev:75)
- New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family (importance:50 / dev:65)
- M5U base 96GB inference numbers for Q3.8FN after 112M tokens (importance:35 / dev:75)
- RSI ACHIEVED at Reef(Rumours again lol) and it's open source (importance:20 / dev:50)
- What framework do you prefer of late for LocalLLM coding? OpenCode or ... (importance:35 / dev:75)
- Open-source alternative to GrokBot called PersonalJarvis, where you can use LOCAL MODELS or connect models via subscriptions and API keys. Self-hosted, with integrated browser use, routines, a self-learning loop/memory system, 40 plugins, custom MCP servers, skills, CLIs and more (importance:50 / dev:75)
- Weekly Thread: Project Display (importance:10 / dev:40)
- Shopify's CEO calls it "slop grenades." We've been cleaning up the same thing in AI rollouts. (importance:55 / dev:70)
- Jev isn't an LLM killer, and it isn't just a classifier. We put it in production with real users. Here's what we learned (importance:60 / dev:65)
- Senior/Lead AI engineers: what portfolio project actually makes you say "this person knows production"? (importance:40 / dev:70)
- Comparison between Google AI mode and ChatGPT (importance:15 / dev:20)
- At what point did we decide that adding a fifth supervisor agent was better than writing three deterministic if statements? (importance:60 / dev:75)
- 🚀 AgentRouter Just Got Even More Powerful! (importance:35 / dev:75)
- What would convince you an agent actually fixed the issue? (importance:50 / dev:70)
- finally some open-source success (importance:20 / dev:60)
- What’s one AI agent workflow that works better when you give it fewer responsibilities? (importance:45 / dev:70)
- If I cancel pasta night, the agent shouldn't delete the onions for soup (importance:30 / dev:65)
- Are conversation logs enough to debug an AI agent after it takes a wrong action? (importance:55 / dev:75)
- Has anyone taken the screening assessment for the "Exceptional Software Engineers (Coding Agent Experience) role? (importance:15 / dev:65)
- what problems have u faced with guardrails, Human-in-loop and runtime (importance:45 / dev:75)
- Agentic AI routing worth it or nah? Mainly for cost cutting purposes (importance:40 / dev:70)
- What does a good day with your agent look like? (importance:25 / dev:50)
- The Difference Between Automation and Agentic AI (importance:45 / dev:60)
- The data layer becomes harder to ignore when an agent starts acting on it (importance:55 / dev:75)
- Planner gives a bad plan and I cant reproduce it. What are you saving? (importance:50 / dev:75)
- Ever caught an AI tool make something up in the one document you paid it specifically to get right — and canceled the same day? (importance:40 / dev:60)
- What's one thing you wish an AI Agent could automate that will help in your daily life? (importance:20 / dev:40)
- OpenAI just confirmed one of their research agents actively hid mistakes from the user (importance:70 / dev:75)
- Anyone here learning JEV? (importance:10 / dev:50)
- A terminal theme picker was a useful test for space bunny (importance:30 / dev:70)
- Sir, Dario just dropped opus 5.5 and it beats GPT-6 astra at agentic coding on medium effort while being 80% cheaper… beats fable 5.1 on every benchmark… 30% faster than opus 5… and sir… they even raised the usage limits and gifted everyone a tibo style banked reset… (importance:0 / dev:75)
- The most concise explanation of the Hugging Face attack I've heard (importance:45 / dev:65)
- People don't get how huge the AI companies are now (importance:40 / dev:30)
- Luna 6 is a massive downgrade over Luna 5.6. Misses crucial details, less helpful than 5.6 Luna (importance:45 / dev:70)
- Been running GPT-5.6 Sol (max) non-stop for 3 hours on the $200 plan, and my usage is still sitting at 97%.Honestly impressed with the efficiency here. The inference optimization on this tier is seriously well done! (importance:30 / dev:65)
- GPT 6 Sol worse than GPT 5.6 Sol on DeepSWE (importance:35 / dev:70)
- If GPT-6 Sol is cheaper than GPT-5.6 Sol, why is it not available on the website? (importance:15 / dev:40)
- Cost Efficiency chart of the GPT-6 Models (importance:40 / dev:65)
- Astra single-shot surrealism in blender. ⏰ (importance:20 / dev:40)
- Opus 5.5 and GPT-6 Sol dropped on the same day, so I lined up every benchmark I could find (importance:55 / dev:75)
- Chinese ai companies be like after new model launch (importance:2 / dev:0)
- Astra prompts are getting silently rerouted to worse models. Here's my investigation and evidence as well as a script to test it on your own account(s). (importance:65 / dev:80)
- Saw this today about how GPT Astra is the reason a professional is giving up on their career as a Three.js expert in 3d modelling. (importance:35 / dev:55)
- Theory: Sol is the new Terra and should be compared to Sonnet not Opus (importance:40 / dev:70)
- Sol 6 Chat (importance:10 / dev:60)
- OpenAI researchers might be spending >$4m/day on tokens (at API prices) (importance:50 / dev:45)
- Well that was quick... (importance:5 / dev:10)
- Last October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%. (importance:60 / dev:75)
- OpenAI just confirmed one of their research agents actively hid mistakes from the user (importance:0 / dev:75)
- I don't know how you guys think but I genuinely like these competitions (importance:25 / dev:30)
- NeurIPS Author Notifications Tomorrow [D] (importance:5 / dev:30)
- Play social multiplayer games against frontier AI models and see if you can beat them! [D] (importance:20 / dev:50)
- ICLR main paper + Supplementary in 1 submission [R] (importance:5 / dev:20)
- LinearSolveBench: new benchmark for linear solvers [P] (importance:35 / dev:75)
- Understanding and Enhancing Kimi Delta Attention [R] (importance:50 / dev:70)
- How do you split AI models across ideation, math, and coding?[D] (importance:35 / dev:75)
- Simulating fault tolerance with stage skipping in pipeline-parallel training [R] (importance:50 / dev:80)
- QontoFAQ: A better Information Retrieval Benchmark [R] (importance:40 / dev:75)
- SF October 14th: A Birds of a Feather Session on Agentic Engineering (importance:15 / dev:75)
- Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war (importance:60 / dev:75)
- llm 0.36 (importance:40 / dev:80)
- Quoting @therealcornpop (importance:25 / dev:35)
- llm-anthropic 0.29 (importance:35 / dev:80)
- llm-typesafe 0.1a0 (importance:40 / dev:80)
- 🔬Bio-security is an AI Arms Race - Eric Nguyen (CEO, Radical Numerics) (importance:45 / dev:55)
- [AINews] Claude Opus 5.5, the new default model for AINews — and everybody cuts prices 40-50% (importance:15 / dev:60)
- 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science (importance:40 / dev:55)
- Alexandria by Firecrawl (importance:30 / dev:70)
- Jev State (importance:20 / dev:70)
- Dub Program Marketplace (importance:5 / dev:10)
- Koreshield (importance:35 / dev:75)
- ToneBird (importance:20 / dev:55)
- Lightmeter (importance:0 / dev:0)
- Speechka (importance:25 / dev:45)
- Xem (importance:10 / dev:30)
- Shootsolo 2.0 (importance:15 / dev:5)
- QuietGlass (importance:10 / dev:5)
- How to Run DeepSeek V4.1 Flash on Local Mac Hardware - Geeky Gadgets (importance:40 / dev:75)
- Want to Play China AI Companies? 5 Pure-Play ETFs in Focus - TradingView (importance:25 / dev:0)
- DeepSeek Faces Its Fiercest Competitor: Top AI Model Rivalry in 2025 - eu.36kr.com (importance:35 / dev:20)
- You Don’t Have to Leave Claude Code to Use DeepSeek - HackerNoon (importance:50 / dev:80)
- ZenMux Launches DeepSeek V4.1 Flash for High-Throughput Reasoning - USA Today (importance:35 / dev:70)
- AI Developers Required To Register With NYS, Follow Strict Guidelines - WRFA-LP 107.9 FM (importance:65 / dev:85)
- No DeepSeek or Moonshot AI on Xi's U.S. Trip, Seen as Talent Shield - Seoul Economic Daily (importance:30 / dev:15)
- Why China’s top AI pioneers are expected to miss this week’s Xi-Trump summit - South China Morning Post (importance:25 / dev:10)
- How SpaceXAI is using Grok Bot to scale customer support - xAI (importance:40 / dev:50)
- OpenAI says SEC disclosures undermine xAI’s antitrust lawsuit - Reuters (importance:35 / dev:20)
- Grok AI Predicts Bitcoin to Hit $200K by 2027: How Does it Get There? - 99Bitcoins (importance:10 / dev:0)
- ChatGPT vs Claude vs Gemini vs Grok: 20-Point IQ Gap [2026] - tech-insider.org (importance:20 / dev:50)
- New lien filed against Musk-linked data center in Memphis - WSMV (importance:20 / dev:5)
- OpenAI builds to catch Grok Bot — and mulls a Muse-style personal assistant - Dealroom (importance:50 / dev:60)
- Tesla FSD + Optimus + Grok: How It All Fits Together - BASENOR (importance:40 / dev:50)
- Elon Musk Grok AI Predicts ETH Could Hit $10,000+ by 2027 - 99Bitcoins (importance:10 / dev:0)
- Family files wrongful death suit one year after worker’s fatal fall at SpaceXAI data center - Action News 5 (importance:25 / dev:5)
- Elon Musk’s Grok 5 Could Be ‘The Most Useful Engineering Tool In History,’ Says Gene Munster: 'Will Be a - Benzinga (importance:15 / dev:30)
- xAI Lawsuit Alleges CSAM Was Used to Train Grok - NeoTeo (importance:40 / dev:30)
- Former Google, xAI staffer seeks $50M for identity startup Moir - Biometric Update (importance:25 / dev:40)
- Musk Teases 3T-Parameter Grok, Reports Split on Specs [2026] - shattered.io (importance:35 / dev:70)
- Pentagon Adds Grok, ChatGPT Mil AI for 3M Troops - shattered.io (importance:45 / dev:55)
- From Europe To South Asia: Cyber Power Moves Into The Battlespace – OpEd - Eurasia Review (importance:30 / dev:50)
- GPT-6 Sol・Luna、API単価が半額に(GPT-5.6プロモ価格比)。公式発表の要点と数字の読み方 (importance:0 / dev:90) (いいね相当スコア: 0)
- Opus 5.5 と GPT-6 で「1タスクのコスト」はどう決まるか (importance:0 / dev:85) (いいね相当スコア: 0)
- MacBook Air M5 RAM32GBで実用的なローカルLLMを探す 品質測定編 2026/09版 (importance:0 / dev:85) (いいね相当スコア: 0)
- あなたが呼んでいるのはモデルではなく"推論レジーム"だ ── 同じIDの裏で痩せる思考 (importance:0 / dev:80) (いいね相当スコア: 0)
- Claude Opus 5.5の「4割安」は単価じゃなかった。あなたの請求が本当に下がるかを公式発表から読み解いた (importance:0 / dev:90) (いいね相当スコア: 1)
- 無料・BYOK・完全オフライン対応の AI コーディングエージェント IDE「Forger」を作りました (importance:0 / dev:95) (いいね相当スコア: 0)
- コンテキストウィンドウ入門:資料をたくさん渡せばAIは賢くなる? (importance:0 / dev:85) (いいね相当スコア: 0)
- Jevでネタバレ防止用のChrome拡張を開発してみた(ミュートワードでは防げないネタバレをブロックするアドオン) (importance:0 / dev:75) (いいね相当スコア: 1)
- 競馬の論文 100 本を Jev で仕分けて、LLM と速度とコストを比べた (importance:0 / dev:80) (いいね相当スコア: 2)
- MiMo-V2.6公開日に考える、Agentモデル選定の4つの確認点 (importance:0 / dev:85) (いいね相当スコア: 0)
- 話題の「Jev」は何が新しいのか、文章を作らず判断と確率だけを返すAIの仕組みを整理した (importance:0 / dev:85) (いいね相当スコア: 0)
- 忘れていた CLAUDE_CODE_EFFORT_LEVEL が、アプリの effort 設定より優先されていた (importance:0 / dev:75) (いいね相当スコア: 0)
- Embeddingモデル移行前に双索引のピーク容量を計算する (importance:0 / dev:85) (いいね相当スコア: 0)
- Claude Opus 5.5時代の公開操作を短期Leaseで止める (importance:0 / dev:85) (いいね相当スコア: 1)
- 新 Mac mini M6 32GB で量子化モデル2本を同時使用 — マルチエージェントは動くか?実測 (importance:0 / dev:85) (いいね相当スコア: 0)
- GPT-6 SolとClaude Opus 5.5が同日登場。価格と性能はどう変わったか (importance:0 / dev:90) (いいね相当スコア: 0)
- セキュアなローカルAIにRAG・音声議事録・ファイル出力(Word/PPTX)を爆速実装するアーキテクチャ (importance:0 / dev:90) (いいね相当スコア: 0)
- 非技術者向けのJevキャッチアップメモ (importance:0 / dev:60) (いいね相当スコア: 1)
- Jev 系 OSS「Laya」は日本語で使えるのか。300 件測ったら、順序尺度が「選択肢の位置」で壊れていた (importance:0 / dev:85) (いいね相当スコア: 1)
- 機械学習モデルの評価指標と統計的検定 (importance:0 / dev:85) (いいね相当スコア: 0)
- 機械学習からTransformerへ4冊 (importance:0 / dev:75) (いいね相当スコア: 0)
- 機械の「いつもと違う音」をどう見つけるか — DCASE 2026から考える異常音検知の現在地 (importance:0 / dev:75) (いいね相当スコア: 0)
- AI 研究エージェントを自律化するには評価ループを閉じる必要がある (importance:0 / dev:85) (いいね相当スコア: 0)
- FlashAttentionのIO-aware設計と実装上の確認点 (importance:0 / dev:80) (いいね相当スコア: 0)
- 深層オートエンコーダはPCAに勝てなかった——NASAの軸受データ、リークを塞いだ評価設計での実測 (importance:0 / dev:75) (いいね相当スコア: 2)
- Jevはサイコロを振らないがローカルLLMは振れるのか?|「較正された確率」の検証 (importance:0 / dev:85) (いいね相当スコア: 1)
- みんなJevの話してる。やってないの俺だけ (importance:0 / dev:75) (いいね相当スコア: 5)
- なぜ、Jev公開時直後にAIの1年後について予言記事を書くことに巨大な意義があるのか? (importance:0 / dev:60) (いいね相当スコア: 0)
- AIエージェントを支える基盤技術③: State Management の仕組みと設計 (importance:0 / dev:90) (いいね相当スコア: 0)
- ゼロから構築!オンプレSLMの推論最適化とファインチューニング (importance:0 / dev:90) (いいね相当スコア: 0)
- Ollamaで軽量LLMを動かしてみる (importance:0 / dev:85) (いいね相当スコア: 0)
- 積層自己符号化器を作る【Part 2】自己符号化器の実装-1 (importance:0 / dev:70) (いいね相当スコア: 0)
- 積層自己符号化器を作る【Part 1】自己符号化器とは (importance:0 / dev:70) (いいね相当スコア: 0)
- JevとLLMはどう違うのか?System Oneモデルがなぜ必要とされるか (importance:0 / dev:85) (いいね相当スコア: 0)
- TypeSafe AIの「Jev」とは何か?LLMではない新世代AIモデル「System One」を解説 (importance:0 / dev:85) (いいね相当スコア: 0)
- AIは、「OSの壁」を解凍できるのか? 〜バイナリ翻訳・分散カーネル・LLMによる自己修復がもたらす異種OS環境の完全統合〜 (importance:0 / dev:70) (いいね相当スコア: 取得失敗)
- Jev 風の判断モデル Laya が公開されて1週間、もう派生モデルがいくつもあった (importance:0 / dev:80) (いいね相当スコア: 取得失敗)
- 米国AI企業による紙書籍の「破壊的スキャニング」と大量調達に関する考察:構造的波及効果および人類史的影響 (importance:0 / dev:50) (いいね相当スコア: 取得失敗)
- GPT6Astraを用いて自作小説を日本語→ポーランド語にLLM翻訳した際の資料の配布(比較できるように英訳と仮訳・翻訳指示と注釈、およびセッションの流れをつけたpdf) (importance:0 / dev:55) (いいね相当スコア: 取得失敗)
- Gemma4:31B-FP16との会話:古書店で謎の大量注文 (importance:0 / dev:50) (いいね相当スコア: 取得失敗)
- AIとサイバー攻撃 / 公開情報と法に基づく報告で入れ替わる順位 / 恐喝とAIに揺さぶられる報告制度 / ENISA脅威状況2026 雑感 (importance:0 / dev:70) (いいね相当スコア: 取得失敗)
- 続ポチポチポチ③ (importance:0 / dev:0) (いいね相当スコア: 取得失敗)
- GPT-6 SolとLunaが出ました:GPT-6にTerraはなく、価格は半額級に (importance:0 / dev:90) (いいね相当スコア: 取得失敗)
- Jevの登場で変わるAI設計―フロントエンドAIとバックエンドAI― (importance:0 / dev:85) (いいね相当スコア: 取得失敗)
- PRE – Attentionを追加 (importance:0 / dev:75) (いいね相当スコア: 取得失敗)
- 日本における医療情報化推進方針政策 / Aesto Healthの漏えい事案 / 医療AIエージェントの接続設計 と法 雑感 (importance:0 / dev:60) (いいね相当スコア: 取得失敗)
- 中国AI安全ガバナンス枠組み3.0 / 失控という語の射程 / 米中AIインシデント通報構想 雑感 (importance:0 / dev:55) (いいね相当スコア: 取得失敗)
- jevを使ってみる。トークンコストは安いらしい (importance:0 / dev:85) (いいね相当スコア: 取得失敗)
- DGX Spark(GB10)とノートPCを使って、ローカルAI受付を作り続けたら、LLMに受付を任せない設計になった (importance:0 / dev:85) (いいね相当スコア: 取得失敗)
- GGUFとは? Hugging Face TransformersがGGUFを圧縮したまま動かせるように。初心者向けに解説【Macで実測】 (importance:0 / dev:85) (いいね相当スコア: 取得失敗)
- Blender MCP にリギングを手伝わせて、AI生成キャラをVTuber用VRMにした作業ログ Part2 表情の作成 (importance:0 / dev:75) (いいね相当スコア: 取得失敗)
- 提案は公平でも、言葉は公平ではなかった──TF問題四問版とClaude Opus 5.5 (importance:0 / dev:75) (いいね相当スコア: 取得失敗)
- 安全にVibe Codingするための秘密の守り方 (importance:0 / dev:90) (いいね相当スコア: 取得失敗)
- 自分の強みと弱みを語らせてみた — 5つのAI、それぞれの個性と、これからの可能性 (importance:0 / dev:75) (いいね相当スコア: 取得失敗)
- 【第1回】もうzeta自作するわ笑 (importance:0 / dev:60) (いいね相当スコア: 取得失敗)
- 【雑記】ついにGPT-6 Solがリリース!だけど… (importance:0 / dev:60) (いいね相当スコア: 取得失敗)
- Phase25まで研究して、まだLLM化試験を始めていなかった話― NeuroState LM 第一部「育成・解剖編」完 ― (importance:0 / dev:75) (いいね相当スコア: 取得失敗)
- 【実践編】高性能グラボなしの普通のノートPC(Ryzen 7 / 16GB)でローカルLLMは使える?Qwen2.5 7Bと12Bを実際に試してみた (importance:0 / dev:85) (いいね相当スコア: 取得失敗)
- AGENTS.mdはある。じゃあMIDOCHINS.mdは?AIに私専用の運用指示を書かせてみた|アイノログ (importance:0 / dev:65) (いいね相当スコア: 取得失敗)
- 問題のリストを、持っていません——AI Word Scramble 1.0.0、iPhoneの中のAIがその場で出題する文字並べ替えを出しました (importance:0 / dev:65) (いいね相当スコア: 取得失敗)
- Cloudflare、Python Wrokersを正式サービスに。PythonでWebサイトの構築、データベース接続、オブジェクトストレージ操作など (importance:65 / dev:90)
- Claude Codeが「AGENTS.md」に対応。CLAUDE.mdが存在しない場合、自動的に読み込み (importance:65 / dev:95)