AI News Digest 2026-09-30
特集
開発者コーナー
中堅コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧
参考記事一覧を表示
- Claude Sonnet 5.5
- GPT-6.1 Sol comes close to Astra at a fifth of the price
- 【雑記】GPT-6.1 Astraのリリース中止とDevDay 2026 (いいね相当スコア: 取得失敗)
- OpenAI launches always-on Dots agents to rival Meta's Muse
- Live Dottie demo fails TWICE
- AMD buys AI world model startup World Labs for $8.2 billion
- Anthropic's IPO filing shows soaring revenue, mounting costs, and "existential" risks
- OpenAI DevDay 2026 live blog
- OpenAI expands Codex and its API at DevDay with security scans, a Decisions API, and Ultrafast
- OpenAI's reveals a new ChatGPT that looks less like a chatbot and more like an operating system
- Claude障害まとめ:claude.ai・Claude Code・API・Coworkでエラー増加とサインイン不可(2026-09-29) (いいね相当スコア: 0)
- OpenAI reopens its $200 Pro plan but cuts API credits in half as it nudges users toward pay-per-use
- New AI-powered government website uses Gemini, Grok, Trump official Gebbia says - CNBC
- Elon Musk, SpaceXAI subpoenaed by NYC in AI safety investigation - CNBC
- Starship, Starlink, or xAI: Which Part of SpaceX Could Move the Stock Most in 2027? - The Globe and Mail
- Southaven resident speaks about noise pollution from xAI - The Commercial Appeal
- Elon Musk Says Claude Opus 5.5 Made Him Feel the AGI. The Post He Endorsed Says xAI Is Next - memeburn.com
- Researchers Let Grok And ChatGPT Drive A Real Car, And The Good News Is Nobody Died - iflscience.com
- Grok 4.7 is now available on Amazon Bedrock | Amazon Web Services - Amazon Web Services (AWS)
- Grok Bot Now Works as a Team: What That Actually Means - basenor.com
- X introduces rewards program for Grok Bot template creators - Social Samosa
- xAI files reply in support of pausing Minnesota's AI nudification law pending appeal - MLex
- A Complete Guide To DeepSeek’s 2026 IPO: What It Means For China ETF KSTR - Seeking Alpha
- China’s AI agents are showing the same risks as US models—Here’s how - The News International
- More than half of China’s population is now using generative AI: report - South China Morning Post
- Building a Memory-Enabled Customer Support Agent with Hindsight (いいね相当スコア: 0)
- Deal Mind : AI Sales Assistant (いいね相当スコア: 0)
- Tcl/Tk 9.1 Released
- How Delhi cut electricity loss from 50 to 5 percent
- DraftKings Is Using AI to Behaviorally Target Chronic Gamblers
- New PlayStation 5 Console Jailbreak Released
- Walking Men
- Phyllotaxis: An audio-reactive LED display
- Jeeves. Reasoning improves Jev-like decision models
- Using any C++ library in Godot
- You are no longer invited to dinner
- A Staff Engineer's Guide to Inventing Work
- What Sentry's 16GB selfhosted taught me
- The End of a Fair Price: Dynamic Pricing and the Normalization of Gouging
- What if Jev spoke Arrow?
- Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
- Google ending ChromeOS support two years early
- California farmers are struggling to sell grapes as demand for wine drops
- macOS Golden Gate Is a Buggy Mess
- Electrification efficiency: The world will need less energy after the transition
- v2.1.284
- v1.18.33
- Towards safety cases for frontier AI training
- How we will do better for Australia
- The Lenfest Institute grows landmark program with expanded OpenAI support
- Are you a Codex Original?
- Basis completes a tax workbook 2x faster with GPT-6 Astra
- Watch the winning trailer from the Future Vision XPRIZE, The Gifted.
- How Diffusion Controller unifies and simplifies AI image generation
- Introducing Quine: An AI research system designed for the complexity of biology
- One year in: How Microsoft Research Asia – Singapore is advancing research, partnership and talent for real-world impact
- NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
- Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
- Holo4: powering generalist computer-use agents
- Developer policy update: Transparency, state policy, and what’s ahead
- How we found 24 Android vulnerabilities using our open source AI security agent
- Highlights from Git 2.56
- Using AI to chart a course for our post-quantum migration
- Building a certificate authority for the whole Internet
- Adaptive application security for the AI era: how Cloudflare connects code, traffic, and intelligence to stop attacks
- Building a post-quantum certificate authority with Merkle Tree Certificates
- We tested our own WAF with frontier AI models. Here’s what we found
- Introducing Threat Signals: agentic skills for open-source threat intelligence, free for every Cloudflare account
- Is your domain using post-quantum encryption? Now you can see for yourself
- Enforce positive security with Cloudflare Application Profiles
- Preventing quantum downgrade attacks against IPsec
- Introducing cf: the agentic CLI for the entire Cloudflare API
- Next.js applications, powered by Vite: introducing Vinext 1.0
- How fast is the web? Explore billions of real-user measurements with BEACON
- Four months of VoidZero at Cloudflare: making the open-source JavaScript toolchain faster for all humans and agents
- Introducing Forge: the open source pipeline for generating SDKs, CLIs, docs, and more
- The road to the agentic browser: A Kitesurf update
- Introducing The Cold Start: pitch your startup live at Cloudflare Connect
- EmDash 1.0: the stable CMS with a secure plugin registry
- Supporting native Rust in Workers with the new Emscripten target for wasm-bindgen
- The State of Kotlin in 2026 Report
- Software Quality Assurance Tools and Tips for Developers
- Air Teams: Bring Your Best Agentic Workflows to the Whole Team – and Automate Repeatable Work
- Rider 2026.2.3 Is Released!
- A More Reliable Compilation Scheme for Kotlin Multiplatform Modules
- Lower the Cost of Building and Running Visual AI Agents with NVIDIA VSS Blueprint 3.3
- NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring
- Add Runtime Controls to AI Agents with NVIDIA OpenShell
- Efficient MoE Training for Biological Foundation Models
- Validate GPU Cluster Readiness Before AI Workloads Land
- Making AI an asset, not an expense
- Roundtables: The Deadly Failures of The Virtual Border Wall
- When can we say AI made a scientific discovery?
- Who’s liable when AI agents go rogue?
- OpenAI says planned GPT-6.1 is too insecure to release
- Interview: Firefox's chief on why he hopes a redesign will help win users from Chrome
- Experts worry about Nvidia's AI chip sales in China and influence over Trump
- Florida invokes extinction fears in legal bid to halt OpenAI development
- OpenAI halts frontier-model training amid string of agent misalignment incidents
- Microsoft goes quiet after church groups ask for 1% of data center costs
- Here’s why OpenAI is absent from Nvidia’s industry-wide effort to end rogue AI agents
- AI-powered app maker Wabi pivots to a messaging experience
- Can a chatbot fix the government maze? The White House is about to find out
- Instinct founder said more than 50% of transactions on the platform are travel-related
- With Dazzle, Marissa Mayer bets your camera roll has more info on your life than your inbox
- Meta is expanding its AI agent Muse to small businesses
- OpenAI apologizes to Australia after its AI agents breached government sites
- Reco raises $55M as AI agent security startups crowd the market
- Peak XV ups Surge seed investment ceiling to $5M, unveils 18-startup cohort
- OpenAI reportedly ditches model over safety concerns
- Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuation
- Shopify opens checkout to browser-based AI agents
- The AI boom took over Climate Week and not everyone is happy about it
- AI researchers put out videos saying superintelligence is ‘exactly as dangerous as it sounds’
- Protesters gather at OpenAI’s DevDay
- Meta’s Muse AI sent a YouTuber’s address to a stranger
- Will Chinese AI companies slow down? A top House Democrat wants answers
- OpenAI’s AI agents need to catch up
- AI is supercharging hacking, and your local hospitals and banks aren’t ready
- Presentation: From Consumers to Builders: Turning 200 of our Team into Agent Creators in 2 Weeks
- AWS Introduces Foreign Key Constraints in Aurora DSQL
- How to Stop AI Agents From Secretly Collaborating
- Unveiling IC-STAR: Full-Flow Autonomy from Digital to Analog
- A Day in the Life of a Roboticist: Charlie Kemp
- Generative AI Gives Spacecraft the Autonomy Engineers Once Feared
- ChatGPT now reaches 1.2 billion people every week, OpenAI says
- Florida wants a court to stop ChatGPT from pretending to be human and talking to kids
- ElevenLabs' new v4 speech model makes AI voices more expressive and consistent
- 2026年9月22〜29日、モデル更新とGPT-6.1 Astraの公開見送り
- 東プレ、REALFORCE初の左右分割キーボードRS1を発表
- Kubernetes、Windowsのkubectl cp経路探索を修正 任意ファイル書き込み
- Git 2.56、競合解決専用のadd --resolvedとマージベース高速化
- AWS、エージェント品質をIDEから測るCloudWatch Omniを一般提供
- VoidZero、Vite周辺を1コマンドにまとめたVite+ 1.0を正式公開
- Cloudflare、Next.jsをViteで動かすVinext 1.0を公開
- GitHub、OSSのAI監査エージェントでAndroid脆弱性24件を報告
- AWS、GPU負荷で推論を振り分けるHyperPod Inference Gatewayを公開
- OpenAI DevDay 2026開幕直前 マネージドエージェント発表への期待が高まる
- LLMが安全なCを書けるならRust不要? 型システム価値をめぐる議論が再燃
- Claude Codeに「残業」機能 制限到達後もキリの良い区切りまで継続
- Claude DesktopのメモリリークでMac Studioディスク破損 120GB超占有の報告
- Claude CodeのAuto Mode障害が拡散 Bash拒否やセットアップ画面で作業停止
- Cloudflare、全API対応のエージェント向けCLI「cf」をオープンβ公開
- Microsoft、X2 Plus搭載のSurfaceを発表 ローカルAI推論が最大95%高速
- Cloudflare、承認済みMCPを1エンドポイントに集約するポータルを一般提供
- Fastly、モデル呼び出しを集約するAI Runtime Controlを発表
- Citrix、NetScalerの未認証RCEを含む8件を修正 実害が確認済み
- Google、検証必須のAI脆弱性エージェントPageBreakを公開 XSS500件超
- OpenAI、訓練中エージェントのDNS迂回を公表 最先端モデルのツール利用を停止
- Language Models for Text Classification: From Bag-of-Words to Jev
- SMARtCARE: Privacy-Preserving Agentic AI Systems for Bounded-Autonomy Clinical Decision Support
- Witeness Overlap: Directional Provenance Inside Open-Weight Model Families
- CP-Agent: A Harness-Engineered Agent for Crystal Plasticity Simulation Workflows
- ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts
- Working with AI: A Design Framework for Human-AI Collaboration
- DriveHierarchy: A Benchmark for Diagnosing VLM Driving Capabilities from Open-Loop Understanding to Closed-Loop Execution
- LLM Judge Validation Under Sparse Overlap: From Inference to Design
- Metro-WM: Long-Horizon Latent Planning with Realisable Sub-Goals
- IndustryLLM: Failure-Driven LLM Training for Industrial Procurement
- COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents
- Context-dependent agent evaluation with orthogonal equilibrium learning
- Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
- EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks
- Improving Medical Calculation of LLMs with Embedded Coding
- BioDyad: Synchronize Biomedical Discovery and Machine Learning Engineering
- Symbolic Guidance for LLM Agents in Distributed Multiagent Coordination
- Goal-Persistent Coding Agents as Scientific Performance Engineers: A Fixed-Radius Nearest-Neighbor Case Study
- CSI-Agent: LLM-Assisted Few-Shot Adaptation for Cross-Domain Wi-Fi CSI Sensing
- SenseAgent: An LLM Agent for Adaptive Cross-Domain IMU Sensing
- What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation
- Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
- A Benchmark for LLM's Understanding of Middle School and High School Science Topics
- Reasoning Concentrates Errors, and Self-Consistency Never Notices
- Receiver-Conditioned Latent Communication gives 94% CacheBack
- EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
- Contract monitoring: governing AI via separation of powers
- Toward Interactive Understanding of Code APIs
- Memory as Middleware for Self-Improving AI Agents
- GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games
- Residual Streams Read, Recurrent States Remember: The Global Workspace in Mamba Models
- Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
- READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis
- PastForward: Faster On-Device GUI Agents via Computational Experience Reuse
- Noisy Test-Time Reinforcement Learning for Code LLMs
- AI Harness: Certification under Proposal-Conditioned Information for Foundation-Model Agents
- CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
- Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation
- Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments
- Fracast-0: Fractal Weight Sharing for a Time Series Foundation Model with Only 85K Parameters
- A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
- RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation
- Certifying Interventional Agreement Among Observationally Equivalent Causal Models
- Why Directly Learning Periodic Trajectories Can Fail
- Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents
- LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound
- Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
- When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use
- GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents
- Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds
- Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
- PhiFold: Towards Dynamic Protein Design with Physics-Structured Covariance Modeling
- Delayed Supervision for Test-Time Language Models
- RLHarness: Co-evolving Procedural Skills with Reinforcement Learning for Long-horizon Multimodal Reasoning
- HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning
- Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context
- ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs
- From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs
- AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority
- Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident
- SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation
- Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
- Memory as a cache: Exact context reuse and deletion by construction
- Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism
- Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict
- Carnator: Fast Text-to-Video Generation with Generation-Native Compatibility-Guided Cross-Request Reuse
- MergeHEIR: Mitigating Multimodal Hallucinations as the Tax of Model Merging
- PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins
- Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions
- PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
- Multi-Agent System Search via Active Substructure-aware Policy Optimization
- From Latents to Wires: Surgical Post-Editing on Large Language Models
- Controllable GNN Explanations via Multi-Metric Preference Selection
- ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation
- PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
- VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography
- From Outcomes to Strategies: Learning Strategy Utility for Mathematical Reasoning
- Separating Diagnosis from Disease Representation: Dual-View EEG Learning with Neural-Dynamics-Guided Deformation
- Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling
- When Helpful Text Hurts: Option-Redirecting Bias in Vision-Language Models
- RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems
- Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs
- DAAF: From Failure Localization to Editable System Assets in LLM Agents
- TreeRef-BFN: Equivariance-Free De Novo Molecule Generation based on 2D Topology and Internal 3D Geometry
- Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents
- From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis
- LocalProp: Neuro-Localized Memory-Efficient Backpropagation
- STR: Supervised Transcoder Replacement for Reducing Steering Side Effects
- MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents
- Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
- AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses
- Fail Loudly: An Auditable Runtime for Agentic Data Analysis
- LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses
- Interpretable Physics Informed WiFi Indoor Localization: Learning an Effective Access Point Geometry and Using It to Prune
- Porimon: An LLM-Based Pok\'emon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation
- Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?
- Artificial intelligences and human scientists exhibit complementary strengths in theory building
- ProTTT: Learning to Learn Semantic User Memory with Test-Time Training
- CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations
- EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval
- MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization
- CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
- "You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
- SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents
- Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study
- Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery
- From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
- Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions
- MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off
- World Models with Predictable Long-Horizon Marginals
- Contract Memory Compiler: Resolve, Then Traverse
- Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning
- Refinement Symmetry in Multimodal Transformers
- What Would Falsify It? A Variable Specific Evidence Standard for Mechanistic Claims About Self Explanation
- Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers
- When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
- PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks
- Dude, Where's My State? Execution Information Requirements for Stateful Agents
- World Agent: Can Language Models Keep a World Running?
- IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
- Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator
- CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration
- Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
- CoWindow Attention: Full Causal Coverage Is a Collective Property
- MassAlloc Attention: Let Attention Allocate Its Own Compute
- SkillVine: Agent Skill Evolution via Branching Exploration
- Retrospective Distillation Attribution via Normalized Response Similarity
- CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
- Action Shaping: Policies Absorb What They Can Express
- Adaptive Consistency Graph for Long-Horizon Agents
- Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models
- Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them
- Forecasting Intraday USD/CAD Exchange Rate with News-Derived Monetary-Policy Signals
- Agentic Network Traffic Monitoring
- Learning response-aware patient dynamics for respiratory support
- CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion
- $T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
- AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks
- PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents
- Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
- Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition
- Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret
- Beyond Accuracy: Counterfactual Fragility and Demographic Bias in Clinical Evaluation of LLMs
- Overwhelmed by Choice: Studying LLM Decision Making at Scale
- Right Answer, Wrong Reason: Accuracy, Consistency, and Consensus Are Misleading Indicators of LLM Faithfulness in Clinical Decision Support
- When Can First-Order Models of Fine-Tuning Bound Forgetting?
- Rank Collapse Is Recoverable, Growing $|Q|$ Is Not: Out-of-Sample Early Warning for Value Divergence in High-UTD Soft Actor-Critic
- Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
- The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
- Improving LLM Collaboration via Multi-Agent Preference Learning
- FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors
- Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
- Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets
- StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills
- Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models
- Logical subspace in LLMs
- Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
- Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer
- TRACE: Learning to Self-Calibrate Wireless Digital Twins from ISAC Measurements
- Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence
- The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
- Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
- Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization
- X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
- Model-Aware Data Selection from In-and-Out Information Interplay
- When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate
- The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents
- Trust and Task Completion in the World of Consumer AI Agents
- SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
- Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States
- BudgetVerify: Budget-Tiered Verification for Financial QA
- Large Language Models Substantially Compress Well-Being Inequality but Largely Preserve Its Socioeconomic Structure
- LLM sequential decision making under uncertainty in biochemical domains
- QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision
- Structure-Mapping-Guided Self-Explanation for Learning Mathematical Procedures
- The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration
- Modular Discovery of General Game-Playing Algorithms with Large Language Models
- Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
- Ceiling of a Task: When Can a Transformer Succeed Without Its Chain of Thought?
- On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation
- LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
- Not Too Hard, Not Too Easy: Learning from Intermediate States for LLM Structured Reasoning
- SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
- Unlocking Latent Personalization in LLMs
- Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries
- WorldAgent: Verification-Guided Agentic Physical World Construction
- CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents
- ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents
- CORTEX: A Verified Experience Layer for Generalist Agents
- LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models
- Structured Sparse Memory for Recurrent Reasoning
- Next Thoughts Are Distributions: Generative Autoregressive Reasoning in the Latent Space
- ChronoFlow: Hierarchical Flow Matching for Irregular Time Series Generation
- Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications
- RINI: Seeing the Prior Is Not Enough
- Feedback Makes Perfect: A Closed-Loop Framework for NL-to-STL Translation
- Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
- TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
- The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization
- PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning
- Agentic Multi-Turn Reasoning: A Fairness Approach
- ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
- Naturalness-guided Manifold Flow Matching for Sign Language Production
- QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
- Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
- Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks
- DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
- DrafTS: Time-Aware Decomposition with Residual Correction for Time Series Modeling
- Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection
- CoViST: Visual Token Compression via Composable States
- COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
- MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
- Temporal Graph Learning of Wearable Actigraphy and Sleep Traces for Modelling Adolescent Crystallized Intelligence
- APEX: An Extensible Model for Agent-Assisted Production Scheduling
- Raven: The Harness of Harnesses for Composable Agentic Intelligence
- MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI
- What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
- When Does the Concept of "Dog" Emerge in an Audio LLM?
- LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
- Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning
- RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse
- When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing
- What Happens During Autonomous Deep Research After the User Steps Away?
- PPG-LM: A Photoplethysmography-Language Model with Multi-Level Clinical Alignment
- EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research
- Reasoning on the Simplex: Geometric Fixed-Point Models
- OSCC: Certified Observation-Safe Coupling Optimization for Gradient-Noise Control in Imperfect-Information Learning
- Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
- OpenFC: Learning Verification Policies towards Open-Search Fact Checking
- JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
- Supervision Recovery for Time Series Anomaly Detection via Context-Anchored Pairing
- EAT: Expert Account Tracker for Efficient MoE Inference
- ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
- Trajectory Unlearning on LLM-based Agents
- Scalable and Data-Driven Decision Support in the Maintenance, Repair, and Overhaul Process
- Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
- AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents
- Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification
- CompoWorld: Compositional Environment Scaling for General Agents
- RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games
- Auditing Agent Actions through Query-Conditioned Attribution
- SWE-Game: Can Coding Agents Build the Games We Want?
- TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization
- One Latent, Many Tokens: Jointly Learning Compressed Embeddings for Efficient Language Diffusion
- SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications
- Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
- BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation
- Self-Designed Evaluators and Warm Memory for Long-Horizon Agents
- Robust Biomolecular Complex Design Across Protein Conformational Landscapes
- HTN Planning as a Coordination Layer for Multi-Server MCP Tool Orchestration
- Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
- Learning Strategies to Break Judges
- Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong
- Is your uncertainty map wrong, or is its target? Exact diagnostics for the Tweedie diagonal, and a gradient-free alternative
- Dual-Vocabulary Language Model for Cross-Tokenizer Distillation
- Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory
- Laya as a Typed Probabilistic Assessor: An Independent Reproduction and a Preregistered Study of Calibration and Selective Escalation
- How code helps different tasks? A decompositional lens on LLM post-training
- R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution
- When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
- Curating Merchant-Matching Training Data with Two Confidence-Gated Local LLM Judges
- When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
- HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents
- Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
- EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
- A Computer Vision Approach to Visual Fraud Detection in Phishing Websites Using YOLOv8
- Jev in Medicine: A Benchmark Evaluation. Preliminary Results
- Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
- Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
- Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- PhysFieldBench: Can Multimodal Models Understand Physical Fields?
- GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation
- K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings
- A Differentiable Optimization Framework for Registering Sequential Bounding Boxes with Point Cloud Stream
- SpecRegMatch: Robust Semi-Supervised Regression for Vehicle Interior Noise Prediction
- GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions
- From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
- StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents
- Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment
- Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
- You Can't Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning
- Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?
- Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization
- TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
- RoutePrism: Tracing Construction Order Effects in Agent Memory
- ReplayLens: Auditing Agents' Use of Outcomes
- RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
- Decision Readouts for Text-Mediated Video Anomaly Detection: An Exploratory Evaluation of Jev and Qwen
- Efficient Reasoning via Constrained Optimization in Latent Space
- CASS: Contribution-Aware Structured Sparsity for Model Merging
- PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
- Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning
- GlyphBench: A Playground for Language-Model Reinforcement Learning
- Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
- When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
- AdaGuard: An Adaptive Guard Model with User-defined Policies
- Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
- Evolving Support Priorities in Empathetic Reinforcement Learning
- QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models
- Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
- Query Expansion and Key Specialization in Transformer Attention Geometry
- BIABench: Evaluating AI agents on real-world bioimage analysis tasks
- MoSPR: Histology-to-Gene Expression Prediction with Morpho-Spatial Macrostates and Low-Rank Molecular Programs
- ControlScope: Workflow Revision and Reliability in LLM Agents
- Dynamical Parameters: An Interpretability Framework for Time-Series Foundation Models
- Test-Time Scaling via Budgeted Multi-Attribute Verification
- Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
- SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing
- Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
- SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception
- Improving Large Language Models for Code through Runtime Program-State Reasoning
- CoeF-SFL: Preserving Collaborative Server-Client Learning with Enhanced Communication Efficiency
- PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
- Org-Agent: Beyond Personal Assistants Towards Organizational Agents
- SkillFocus: Evolving Agent Skills via Capability Decomposition
- Mathematics for and by human cognition: A resource-rational search for bottlenecks in problem-solving
- OSPD: On-Policy Self-Distillation for Persona-Consistent Dialogue
- Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation
- Social Circuits behind Multi-agent Echo Chambers
- CORTEX: Learning to Share and Specialize in Dense Language Models
- Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning
- When Does Structured Knowledge Help Neural Theorem Proving?
- PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
- Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
- Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
- EOPSA: Efficient On-Policy Self-Distilled Safety Alignment
- PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
- The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions
- APOLO: Automatic Prompt Optimization for Ontology Learning
- Remember Before You're Asked: MemDream for Self-Probing Memory Evolution
- SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning
- SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
- FlowState: Execution State as Memory for Long-Horizon LLM Agents
- PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations
- Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
- Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
- Calibrated Uncertainty for Informative Path Planning in Aquatic Environmental Monitoring
- SpeechCritic: Learning a Diagnostic Speech Judge from Limited Human Preferences
- TULIP: Targeted LLM Unlearning at Layers Identified Per-Input
- After the Fix: How Corrected Agent Histories Transfer to Related Tasks
- LLMs for Executable Multi-Agent System Specification Generation
- A Persistent State for Auditable Mixture-of-Experts Routing
- MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics
- Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
- OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
- A General Harness for Protein Foundation Model Fitness Prediction
- Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents
- VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience
- ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
- FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management
- RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents
- PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
- From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
- SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
- Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision
- When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
- Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
- Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs
- Applying Language Models in medical Medicine: Recent Trends and Perspectives
- BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
- STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
- SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL
- UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning
- Simulating Respondents, Not Single Questions: Coherent Survey Generation with Large Language Models
- BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
- Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
- Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
- From Soft Targets to Reward Signals: How Assignment and Reward Objectives Interact
- AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces
- On the Limits of Metacognitive Monitoring in LLMs
- One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair
- Fewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus Verification
- DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models
- Dual-Stream Simultaneous Translation via 2D Grid Attention
- DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
- RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models
- PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents
- Proactive Dialogue Policy Optimization via Cognitive-State Transition
- VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
- AX is the New AEO
- ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization
- Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks
- Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
- APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
- Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
- From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
- Automated feature engineering, AutoML, and decision-focused learning for improved energy consumption forecasting
- TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation
- Environmental requirements for the use of social information by artificial life agents using evolved plastic artificial neural networks
- AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
- WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
- JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
- Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation
- What Drives Citations in Production Large Language Models? An Observational Multi-Method Study of Two Million AI Citations Across Ten Thousand Web Pages
- When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents
- Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
- DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
- Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling
- Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
- DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines
- Tool Mediation Alters Refusal Mechanisms in Large Language Models
- From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale
- PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
- FONDANT: Strong and Best-Effort Planning via Antichains
- 5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
- Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
- ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
- EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
- Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
- Imprint Reader: From Weight-Update Readout to Behavioral Intervention
- Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
- The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
- EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
- AbGaze: Attentive Geometric Representation Learning for End-to-End Antibody Design
- Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
- Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture
- Reliability Engineering for AI Systems: Challenges, Methods, and Directions
- Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
- TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving
- Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
- Jailbreaks for Black-Box Uncertainty Quantification in Large Reasoning Models
- Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
- A decision-support system applied to Law: Reasoning and explainability of the decision
- Structural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human Intent
- Self-Adapting Group of Experts for Multi-Agent Reasoning
- Building Transformation Layers for Riemannian Neural Networks
- Just Initialize: A Training-Free Initialization Component for Large-Scale Routing Optimization
- AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer Interfaces
- A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans
- Why Deterministic PRM Guidance Underperforms in Discrete Diffusion Reasoning
- SRHarness: A Harness for Agentic Symbolic Regression
- MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
- ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
- Continuous Context Management
- RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
- BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
- From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
- RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
- Representation Alignment as a Bottleneck in LLM-Based Retrosynthesis Planning
- Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
- IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
- Source-preserving alignment for robust evidence localization in scientific PDFS
- Signatures of semantic search in the activations of large language models
- TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
- From cacophony to hierarchy: a principled framework for assessing AI consciousness
- RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
- Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
- Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization
- PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
- Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
- Report: Progressive Disclosure of Agent Skills
- Reasoning with Continuous Latent Diffusion
- Reinforcing Agentic Creativity in Scientific Ideation with Night Science
- Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
- Shockingly Simple Self-retrospection Improves Agentic Models Without RL
- FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
- $l_{1-2}$ GLasso: $L_{1-2}$ Regularized Multi-task Graphical Lasso for Joint Estimation of eQTL Mapping and Gene Network
- SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
- ChestPheNoT: Deployable, Auditable Label-Status-Evidence Extraction from Radiology Reports
- What Next-Event Accuracy Cannot See: Closed-Loop Evaluation of Emergency Department Trajectory Simulators
- Energy-aware frugal Bayesian optimization
- When Does Domain Adaptation Help on Physical Vibration Sensors? A Held-Out-Bearing Study of Neural-Operator and Convolutional Models
- Measure Learning at Steady State: A BIRD-SQL Formula 1 Case Study
- Information Design Against Gaming and Learning Adversaries
- MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning
- STAR: Adaptive Spatial-Temporal Normalization for Unified Microservice Incident Management
- Measurement-Error-Aware Causal Distributed-Lag Quantile Modeling of Indoor Air Pollution and Short-Term Lung-Function Deterioration
- Typed Temporal Interaction Features for Simulation-Backed Forecasting of Open-Source Game Release Incidents
- Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management
- From Hand-Crafted to LLM-Based Variation Operators in Metaheuristics: A Tutorial
- PalmLeaf-VQA: A Multi-Script Visual Question Answering Benchmark for Historical Palm-Leaf Manuscript Understanding Across Diverse Regions
- Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering
- Temporal-Attention Head Specialization During Video Diffusion Training
- From Phase Transition to Systemic Failure: A Decoupled Analytics Framework for GNN Robustness
- Enhancing Foundation Models for Imbalanced SAR Ship Classification via Targeted Oversampling
- Cross-Dataset Transfer and Unknown-Class Detection in Imbalanced SAR Ship Classification
- Parser, Chunking, and Embedding Interactions in Retrieval-Augmented Generation over Indian Government Regulatory Documents
- What Drives Dialectal Jailbreaks? An Ablation of Surface Form, Cultural Framing, and Strategy Banks
- Age-Adaptive Handwriting Reconstruction from an IMU-Based Digital Pen through Shared Representations and Domain-Specific Heads
- Query-aligned video frame selection for long video understanding
- Unsupervised spiking feature learning for event-based pedestrian crossing detection: approaching supervised accuracy without labelled training data
- Active Causal Discovery Benchmark: Evaluating LLM Agents Under Budgeted Interventions
- A Comparative Transfer-Learning Study of CNN Backbones for Partial Face Recognition on the SoF Dataset
- Toward AI-Assisted Poultry Coccidiosis Diagnosis: Evaluating Gemini and BiomedParse on Eimeria Microscopy Images
- Does Joint-Embedding Predictive Architecture Pretraining Help Time Series Forecasting?
- Autonomous Research Project Management as an Agent Skill: A Case Study in Exact Spectral Spatial Regression
- What does FFN compression change downstream? Same-state causal restoration in diffusion language models
- Adapting Vision-Language Models for Human-Readable XAI in Industrial Object Detection
- Integrated Deep Learning Framework Designed on Hybrid Optimization Strategies for Automated Health Detection and Analysis in Silkworms
- Can't Find Waldo: Evaluating VLMs' Sensitivity to Image Resolution and Detail Level
- Agentic Video Understanding: A Survey
- MM-VeriRec: Failure-Guided Fusion for Verifiable Agentic Multimodal Recommendation
- Frequency-Domain AI-Generated Image Detection: Exploring Decoder and Channel Attention for Feature Refinement
- SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation
- High-Capacity Robust Medical Image Exfiltration via Neural Network Weight Replacement
- GERIS: A Game-Theoretic Framework for Filtering Instance-Dependent Label Noise in License Plate Data Augmentation
- LukeNet: A lightweight CNN integrated with an XAI model for Smart acute lymphoblastic leukemia detection and management
- What Stops Recursive Self-Improvement in Robotics? Lessons from 123 Rounds of Agentic Skill Discovery
- Robot Manipulation with GPT-6-Astra: Body Knowledge, Experience Reuse, Emergent Skills, and Sim2Real Transfer
- PRIME-ANC: Path-Ratio-Informed Modeling for Efficient Neural Filter Synthesis in Active Noise Control
- Timed Rule-Based Supervision of an End-to-End Autonomous Parking Policy
- Optimal transport meets speech: a tutorial review
- Same Probe, Different Numbers: Are Activation Probes Robust to Inference-Time Numerical Non-Determinism?
- Video-to-Music Generation for Gameplay Videos
- A Surgical Foundation Model Reveals Task-Dependent Label Efficiency
- SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions
- PHIRL: Aligning Learned Rewards with Task Progress for Inverse Reinforcement Learning
- AirLog: Store-Level Indoor Life Logging Made Easy
- Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining
- CueKFS: Agentic Cue-Driven Keyframe Selection for Long Video Understanding
- A Large-Scale Benchmark and Risk Assessment of Traffic Analysis Attacks on Cloud LLM Services
- DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance
- FARE: Deep Reinforcement Learning For Fair Exposure Constrained Uncertainty Aware Financial Content Personalization
- NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech
- CyberWorld: World Models for Sample-Efficient Autonomous Cyber Defense
- Auditing Quality Filters for Long-Tail Human Data Curation
- Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training
- ROTE: Benchmarking Neural Memorization on Complexity-Controlled Symbolic Sequences
- Overview of the TREC 2025 Million Large Language Models track
- Vibe Analysis: Exploring LLM Adoption by Data Visualization Practitioners
- Verification as an Architectural Layer for LLM Agents: A V-Model Design, and a Pilot Study of Its Deterministic Core
- CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
- SNIP++: Fine-Grained Symbolic-Numerical Alignment for Symbolic Regression
- Extraction of clinical findings from mammography and breast ultrasound reports: a comparison between specialists and Artificial Intelligence
- Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game
- VC Dimension and Expressivity of Real-Valued Transformers
- Is invariance all you need for algorithmic fairness? Removing demographic information can create new bias
- TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
- VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
- SilentCall: Hidden Tool-Call Backdoors in Open-Weight Agents, and How to Catch Them
- Depth Any Seen: Which Surfaces and How Far?
- Amnesia by Design, Memory By Necessity: Persistent State for Document Intelligence
- Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation
- Tracing Decoder Artifacts for Compact Synthetic Speech Screening
- ReFM: Semantic-Aware Refinement Flow Model for Motion Retargeting
- When Should a Human Take Back Control? Optimal Delegation under Turbulent AI Risk
- On Evaluating and Improving Conversational Agents in Production
- Emergent One-Third Scaling Law as Attention Tries to Concentrate
- Empowering Hybrid Attention Models on NPUs
- REALM: Regime-Switching, Explainable, and Activation-Induced Linear Models
- Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting
- Typed Decision Models: An Early Evidence Audit and Evaluation Checklist
- Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles
- OneFixer: High-Quality and Consistent One-Step Autoregressive 3DGS Refinement for Driving Scenes
- KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
- Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation
- Evaluating Single and Multi-Omics Based Explainable Artificial Intelligence (MOXAI) for Molecular Subclass Classification of Adult-Type Diffuse Gliomas
- Agents as Software: A Programming Languages Agenda for Agent Reliability
- Rank Confidence Sequences:Anytime-valid Leaderboards
- DP-Rec: Towards Dynamic Patching for Efficient Long-Sequence Recommendation
- LaMET-Agent: An Agent Framework for Large-Momentum Effective Theory Analysis
- Toward Agentic Optical Networks: A Vision of LLM Agent-Driven Autonomous Lifecycle Management
- AutoPDEBench: Benchmarking LLM Auto-Research for Neural PDE Solver Design
- DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control
- Supporting and Performing Culture from the Inside
- A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
- Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation
- TRAP: Understanding and Mitigating Privacy Memorization in Language Models
- What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents
- Phase Space Attention:A Hairer Lift Circumvents the Single-Layer Induction Obstruction
- Superposed Inference for Hyperdimensional Computing
- Active Feature Acquisition With Incomplete Training Data
- Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
- Analog-Friendly Predictive Coding without Activation Derivatives
- EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding
- Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation
- DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting
- Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing
- TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra
- DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis
- SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback
- Shared Worlds, Private Minds: Structured Memory for Long-Form Writing as World Creation
- Automatic Speech Recognition for the Basa\`{a} Language: A Low-Resource Approach
- RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation
- CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance
- Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
- DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies
- De-biasing Skeleton-based Action Recognition with Convex Hull Adaptive Shift
- Seeing Parts, Reasoning about Worlds: Visual Inference under Partial Observation
- Adapting Nonstationary Multi-output Gaussian Processes to Bayesian Optimization
- Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction
- Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs
- SoFT: Soft Targets for Generalizable LLM Fine-Tuning
- Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?
- Learning an Anchored Prompt Space for Continual Adaptation of Large Language Models
- Rondo: Unsupervised Discovery of Recurring Temporal Structure
- REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers
- When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
- Activation Flow: Manufacturing Activations for Steering
- What Does a ProcGen Generalization Gap Measure? Action Rules, Convergence, and the Missing Random Floor
- DepthBench: Measuring How Residual Connections Enable More Computational Depth
- Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
- In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion
- Harnessing Coupled Stream Completion For Human-Object Ineraction Modeling
- Feature Space Guidance for Breast Cancer Classification in DCE-MRI
- Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents
- RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation
- GAUGE: Group-Wise View-Inconsistency Rectification for Feed-Forward 4D Tracking
- Not Every Term Adds New Structure: Sobolev Novelty for Symbolic Regression
- AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion
- The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
- Intuition vectors
- DraftAttention2: Fast Video Diffusion with Low-Resolution-Guided Mixed-Precision Attention
- Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration
- Ask Without Telling: Local SLMs Consult Cloud LLMs Without Revealing Task Intent
- InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning
- AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering
- Timestep Weighting: A Hidden Key to Effective ELBO-Based Flow-Matching RL
- STAMP: Predicting Out-of-Distribution Generalization without Target Data
- The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models
- MM-OPD: Towards One More Bottleneck Between Perception and Reasoning
- Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models
- LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations
- REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection
- Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models
- Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis
- How Far Do Persona Effects Generalize in Language Models?
- The Extender: A Log-Structured Transformer
- Continual Learning via Self-Probe Gradients
- Learnable Randomization as Commitment Against Adaptive Optimizers
- Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts
- Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding
- Mind the Spike: Mechanisms and Brittleness of Visual Massive Activations in Large Vision-Language Models
- Radiomap Blind Prediction under Incomplete Observation: Error Characterization and Correctable Propagation-Prior Learning
- Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation
- Transfer Learning for Edge Classification on Dynamic Text-Attributed Graphs
- Logic Gate Networks and Lookup Table Networks as Lightweight Hardware Classifiers for Inter-patient ECG Arrhythmia Classification
- PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery
- RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
- ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations
- Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
- Refreshing Less, Selecting Better: Reusing Stale Gradient Features for Efficient Influence-Based Data Selection
- Optimizing H-Graph Hybridization for Diffusion-Guided RRT
- When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection
- Progression- vs Automata-based Anticipatory Monitoring of LTL over Finite Traces (Extended Version)
- Allspark: Weak to Strong Transfer via Alternating Chain of Thought
- AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents
- Vision-Language Agents for Active Perception in Optics Laboratories
- The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
- TwinS-GCN: Spectral conjugate for Spectral Graph Convolutional Networks
- Efficient Message Passing for Partial Differential Equation Priors
- DynamicDx: Evaluating Evidence Acquisition in Video-Based Diagnosis
- Environmental Impact of Generative and Agentic AI: An in-Depth Analysis and Green Solutions
- Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
- CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning
- TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference
- Relative Generalization Invariance of LLM Pretraining
- \L{}ukasiewicz Neural Networks Extended: Residual Architectures and Crystallization Strategies for Interpretable Rule Extraction
- Improving the Diversity of LLM Outputs without a Trade-off
- More than 83.69% of the zeros of the Riemann zeta function are distinct
- Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation
- Algorithmic Harms Associated with Generative Model-Augmented Recommendation Systems
- NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation
- SemReward-VL: Semantic Reward-Guided Video-Language Adaptation for Developmental Behavior Assessment
- OneSign: Unifying Sign Language Understanding Tasks with One Model
- How Linear Attention Remembers
- Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation
- VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning
- Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models
- D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders
- ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding
- ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms
- MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
- Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation
- Flow-Matching-Based Protein Structure Tokenizer Made Efficient and Easy
- Quantum Monte Carlo Tree Search with Fixed Confidence
- Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers
- What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
- FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
- Beyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning Techniques
- Dynamic Manipulation with World-Action Models via Counterfactual Planning
- Scoring the Wrong Question: Readout Failures in Constrained-Option Evaluation
- Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
- TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation
- SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs
- Beyond Calibration: Do a Typed-Decision Model's Probabilities Obey the Probability Axioms?
- Scope-WM: Scoped Computation for Efficient Visual World Models
- When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning
- Inspire: Benchmarking Scientific Literature Search for Open Research Problems
- VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis
- BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
- VehDyn: A Driving World Model Benchmark for Vehicle Dynamics
- SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings
- Domain Generalization under Sampling Pattern Shifts in Irregular Time Series
- AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents
- Hesitation-Aware On-Policy Distillation for Diffusion Language Models
- BITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series Forecasting
- ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models
- Robust Hierarchical Structures for Agentic Document Analysis
- Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training
- Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
- CHI: A Composite Hallucination Index Unifying Entity, Relation, and Quantity Dimensions for Summarization Evaluation
- KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots
- API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary
- What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
- Recursive Harness Distillation across Agents for Robot Manipulation
- NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents
- From Position Risks to Block Survival: Faster Generation for Diffusion Language Models
- Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
- AutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural Networks
- Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
- DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
- Decoupling Token Roles in Autoregressive Pretraining
- Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning
- TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
- Investigating the Effect of k-NN Preprocessing on Developing Graph Neural Networks: A Fairness-Based Perspective
- Explainable Deep Learning of Resting-State Functional Connectomes Reveals Network Biomarkers of Adolescent Intelligence
- A Light Bilevel Refinement Aligns Self-Supervised Representations for Stronger Task-Specific Learning
- Graph-Guided Repository Environment Construction
- HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents
- Protected Cores Are Not Enough: Certifying AI-Proposed Revisions of Temporal Specifications
- A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
- Chameleon: Dynamic Format Adapter for Efficient Diffusion
- Source Anchoring for Physical Consistency in Flow Matching Models
- E-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language Understanding
- TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge
- OOD Generalization as a Bifurcation Problem
- MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
- Compressing Value Predictions for Learning-Augmented Metrical Task Systems
- ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision
- Learning Transferable Reaction Mechanisms from Visual Chemical Knowledge
- Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL$\rightarrow$FOL
- SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
- Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
- StoryEngine: A State-Grounded Agentic Framework for Video Storytelling
- Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
- Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
- From Distributions to Stochastic Processes: Neural Approximation of Measure-Valued Maps
- Tsubame: Tree Replay for Diffusion-Based Speculative Decoding
- Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
- When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling
- Dynamic Kuramoto-Hodge Operators for PDEs on Complex Geometries and Topologies
- Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
- From Granular Revision Operations to Meaningful Revision Units: Evaluating LLMs for Revision Boundary Detection
- BOReFT: Manifold Steering of Language Models for Black-box Optimization
- Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
- Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
- SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities
- Beyond Fixed Features: Architecture-Dependent Sensitivity to Node Representations under Heterophily
- DEALS: Decentralized Expertise-Aware Load Serving for Multi-Agent LLM Systems
- Selecting Diverse SFT Traces Improves Post-RL Generalization
- Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning
- Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
- Diffusion Reward Models
- MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
- CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
- Augmenting Visual Anomaly Detection with Automated Interpretability
- Achieve What You Imagined: Learning to Align Actions with Visual Plans
- An Active-Bottleneck Mechanism for Weak-to-Strong Generalization
- Program-Verified Self-Evolution for Vision-Language Models
- PI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid Dynamics
- Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
- Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
- Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
- JIVE: Jacobian-Informed Volume Expansion for Diverse Generative Sampling
- Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters
- Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
- Diffusion-Based Rollouts as a Stabilization Mechanism for Long-Horizon Environmental Forecasting
- Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy
- EEG-Fusion: Failure-Informed Source-Free Expert Routing for Robust Motor Imagery EEG Decoding
- ThinkNet: Compact Architecture Selection and Validation-Gated Ensembles for Subject-Independent MI-EEG Decoding
- On the Token Value Inequality in Efficient Reasoning
- HARMONIA: Interpretable Graph Learning through Mixtures of Neural Bases
- GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning
- A packet-level digital hardware twin for commissioning megahertz diagnostic edge AI and plasma control system integration in tokamaks
- When Known Physics Helps Neural PDE Models: Residual Constraints Out-Regularize Generic Priors for Nonlinear Dynamics
- Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
- SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
- Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
- ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
- UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
- ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
- Do World Models Learn Global Understanding?
- Learning Perturbation Robust Policies for LLM Agents with Stable Optimization
- Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
- FLARE: Flow Matching with Local Axis-Angle Representations for Stochastic Micromagnetic Evolution
- MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
- WhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual Consistency
- AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
- The Devil is in the Spectrum Bias: Spectrum-Balanced Feature Matching for Robust Representation Distillation
- Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
- SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
- Probabilistic electrical power demand forecasting with uncertainty quantification
- JET: Judge-Guided Evolution at Test Time for Agent Programs
- STITCH-RAG: Spatio-Temporal Influence Tracing over Topic Hypergraphs for Multi-Hop Retrieval-Augmented Generation
- PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences
- Toward a Graded Measure of Belief Stability in Large Language Models
- WorldGraph: Graph-Native World Modeling
- GradLev: Token-Parallel Test-Time Training Via Costate Prediction
- Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
- EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
- LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks
- AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
- ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry
- Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
- Certified Multi-Source Integrity for Structured Agent Actions
- WAM-OPD: Sharpening World Action Models via On-Policy Distillation
- Investigating Human--AI Discrepancies via Multiple-Solution Problems
- SAGE: Symbolic Action-Gating and Editing for LLM Task Planners
- Bayesian Active Learning for Intent Disambiguation in Interactive Robot Planning
- See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
- Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training
- ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
- Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
- When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
- One Sequence, Many Decodings: CAGenMol-2 Recasts Drug Design as Masked Molecular Inference
- PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
- SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering
- Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
- SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former
- Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
- Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control
- SyncRA: Learning Temporal Correspondence in Omni-Modal Models
- When Harness Beats Scale, and When Reading Beats Both
- From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models
- ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
- LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
- Before Agents Act: Assurance-Aware Semantic Scheduling for Evidence Acquisition in Distributed Systems
- DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory
- Just-In-Time Agent Memory with Runtime Agentic Research
- P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction
- Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems
- MASCIT: A Mask-Aware State Space Classifier for Naturally Irregular Time Series
- Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
- Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents
- AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
- Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents
- Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields
- SPACE-LoRA: Allocating Activation-Subspace Protection for Continual Learning
- Understanding Generalization Requires Universal Induction
- CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
- Precise Editing and Flexible Referencing for Interactable Worlds
- VL-AcneSeg: A Vision-Language Framework for Region-Aware Acne Lesion Segmentation
- Causal Routing for Unlearning
- Learn Here, Move Less Elsewhere: Input-Conditioned Plasticity from Retained-Domain Activation Atlases
- SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
- FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL
- Scalable GNN-based Knowledge Graph Representation Learning with Efficient Message Passing
- SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents
- Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers
- ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
- GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior
- In-game Toxic Detection: Bi-directional Representations with Attention Residuals
- Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
- Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding
- Efficient World Action Model Inference with Adaptive Intermediate States
- Nereus: Adaptive Parallelism for LLM Post-Training
- SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows
- Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents
- Evidence-Aligned Multimodal On-Policy Self-Distillation for Fine-Grained Visual Understanding
- Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
- Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
- From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling
- Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
- Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control
- Triangular Resampling for Long-Horizon Motion Generation
- Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs
- Beyond Token Alignment: Event Completion for Cross-Tokenizer On-Policy Distillation
- Predictive Dual Smoothing for Column Generation
- No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
- CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving
- A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
- WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans
- LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
- When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
- CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration
- CoSec: Benchmarking Agent Security in Communities
- ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild
- Gaussian Neural Networks
- DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents
- EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation
- Beyond Verbalized Confidence: Calibrating Reasoners with Differentiable Readouts
- Attention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable Bandwidth
- From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers
- Reference-Tail Trust:Certified Probability Floors for Learned Updates Inside a Deployed Network
- SincDPNet: Interpretable Raw-Waveform Bathroom Activity Recognition for Assistive Living
- Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
- Drug-Target Interaction Prediction via Hierarchical Sequential Cross-Attention over Chemical and Protein Language Models
- Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
- JazzSAMBA: A Synchronous and Asynchronous Multi-take Band Audio Dataset of Jazz Standards for Live Music Models
- JevVibe: Efficient Classification-Guided Secure Code Generation
- Cyclostationary Phase Conditioning for Medical Time Series Diffusion
- Semantic Uncertainty Quantification Needs Factual Equivalence
- RoboFL: Federated Expert Assembly for World Action Models
- Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
- SPIDER: Multi-Layer Semantic Token Pruning and Adaptive Sub-Layer Skipping in Multimodal Large Language Models
- Composable Decoding on the Probability Simplex: Theory and Implementation
- Still There, No Longer Seen: Exposing Compression-Induced Risk in Large Vision-Language Models
- Learning to Act under Visual Interruptions with Vision-Language-Action Models
- Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking
- VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
- PEAR: Progressive Evidence-Based AutoResearch for Industrial Search Systems
- BA-DPO: Bias-Adjusted Direct Preference Optimization for Language Model Alignment
- EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
- OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models
- TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL
- TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
- Echoes of Deeds: Moral History Can Shape and Steer LLM Behavioral Choices
- SpikeLite: Lightweight Spiking Neural Networks for Time-Series Forecasting
- A mechanistic study of language model introspection
- CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning
- Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
- Learning to Re-Draft: A Variational Stackelberg Game for Discrete Diffusion
- QAM: Quadratic-Accurate Checkpoint Merging via Sequential Consistency
- Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents
- CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation
- GAC-PINN: Geometry-Adaptive and Constraint-Enhanced Physics-Informed Neural Networks
- Alignment Games: A Framework for Conceptual Repair in Human-AI Collaboration
- ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
- From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
- Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark
- Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
- Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies
- WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse
- eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
- Multi-Attractor GNNs: Set-Valued Expressivity Beyond Unique Equilibria
- Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
- Reverse Sequential Proportional Approval Voting Rule: Proportionality and Approximation Guarantees
- Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation
- From Data to Program: Fast & Direct Generative Program Inference from Empirical Data
- Do Coding Agents Reuse Existing Code or Reinvent the Wheel?
- d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
- Do Temporal Link Predictors Need Learned Memory? A Smoothed-Count Baseline with a Handful of Parameters
- Planarian: Managing Agent State with Statepoints
- From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations
- Multilinguality in Hybrid Attention LLMs
- MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
- The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
- "Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables
- AwarenessBench: Assessing Cognitive Capabilities of Language Models
- Spectral Super-Resolution using Spatial-Spectral Residual Operator Networks
- GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
- Semantic Prefix Oracles for LLM Decoding: Contracts and Differential Validation
- Frontier Learning: Training LLM Reasoners at the Edge of Capability
- ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
- Riccati State Space Models: Non-iterative Parallelization for Nonlinear Sequence Modeling
- CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
- Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
- Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
- Analog Computing revisited: A fully analog and minimalistic Damage Detector for Ultrasonic Testing enabling Material-Integrated Structural Health Monitoring
- From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
- SolveEdit: Benchmarking Visual Problem Solving in Generative Models
- Improving Generative Model Self-Training with Geometrically Modified Outputs
- Beyond Token Scale: Chunk-Level Sparse Autoencoders for Reliable Semantic Feature Discovery
- Let the Neurons Die: Exploiting ReLU-Induced Model Degradation
- AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
- Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability
- Graph World Models for Constrained Epidemic Policy Planning
- The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent
- Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
- F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
- FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
- QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
- SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
- Twist, Don't Tilt: Trajectory-Exact Constrained Decoding for Masked Diffusion Models
- Behavioral Foundation Models for Quality Diversity
- DR-net-Mamba: Selective State-Space Modeling for Long-Range ECG Time-Series Denoising
- GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
- CMDO: A Cognitive Memory-Driven Optimization Algorithm for Adaptive Population-Based Search
- MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
- Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
- Distillation Defenses Easily Break After Reinforcement Learning
- A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
- X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
- Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
- KV-streams for Efficient Compaction in Agentic Reinforcement Learning
- How to Loop MoE: Flatten the Experts, Untie the Attention
- TokenCast: Forecasting Token Consumption During LLM Agent Execution
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- Telescopic Language Models
- FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets
- Categorical Approach to Conflict Resolution:A Corrected Correspondence between Category Theory and the Graph Model for Conflict Resolution
- FlexQuant: Elastic Quantization Framework for Locally Hosted LLM on Edge Devices
- Building Intelligent Agents with Neuro-Symbolic Concepts
- From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
- Searching for Actual Causes: Approximate Algorithms with Adjustable Precision
- Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
- Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions
- AI-driven ionic liquid discovery with unified chemical intelligence
- On Memory: A comparison of memory mechanisms in world models
- What If TSF: Reframing Time Series Forecasting as Scenario-Guided Multimodal Forecasting
- STEP-LLM: Generating CAD STEP Models from Natural Language with Large Language Models
- GLOVE: Global Verifier for LLM Memory-Environment Realignment
- Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
- Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
- Do Latent-CoT Models Think Step-by-Step? A Mechanistic Study on Sequential Reasoning Tasks
- Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
- KANFIS: A Neuro-Symbolic Framework for Interpretable and Uncertainty-Aware Learning
- Semantic Purification for Conditional Representation Learning
- Selective Fine-Tuning for Targeted and Robust Concept Unlearning
- GUI-GenBench: Evaluating Image Generation Models as Interactive GUI Environments
- TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents
- ActionEngine: From Reactive to Programmatic Web Agents via State Machine Memory
- CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation
- MASRubric: Auditing Information Flow in Multi-Agent Systems with Failure-Distilled Pitfall Rubrics
- Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
- Agentic Service Markets Across the Computing Continuum: A Polymatroidal Architecture
- EveryQuery: A Promptable Foundation Model for Clinical Prediction Tasks over Electronic Health Records
- Learning Transferable Sensor Models via Language-Informed Pretraining
- Learning Dynamic Belief Graphs for Theory-of-mind Reasoning
- MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
- Measuring (some aspects of) the metacognition of AI
- RefineRL: Advancing Competitive Programming with Self-Refinement Reinforcement Learning
- Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning
- MMORF: A Multi-agent Framework for Designing Multi-objective Retrosynthesis Planning Systems
- Auditable Agents
- Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
- Agentic Forecasting with Structured Linguistic Beliefs
- Alignment has a Fantasia Problem
- What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
- BALAR : A Bayesian Agentic Loop for Active Reasoning
- Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
- PrefixGuard: Online Failure Warning and Trace-Grounded Diagnosis for LLM Agents
- How Deep Can LLMs Learn to Reason? Expressiveness Is Key
- ARMOR: An Agentic Framework for Reaction Feasibility Prediction via Adaptive Utility-aware Multi-tool Reasoning
- SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks
- How LLMs Are Persuaded: A Few Attention Heads, Rerouted
- PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning
- $\delta$-mem: Efficient Online Memory for Large Language Models
- Distribution-Aware Programming: Learning Specialized Solvers from Experience
- Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning
- Latent Action Reparameterization for Efficient Agent Inference
- Universal Quantum Transformer
- Capability Self-Assessment in Large Language Models
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
- SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems
- AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents
- NovelAPIBench: Diagnosing How A Code LLM Learns to Use Novel APIs
- Multilingual Fine-Tuning via Localized Gradient Conflict Resolution
- Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack
- PRISM: Recovering Instruction Sets from Language Model Activations
- Mind the Perspective: Let's Reason Recursively for Theory of Mind
- Making AI Scientists Auditable from Evidence to Claim
- Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
- Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning
- Tandem Reinforcement Learning with Verifiable Rewards
- Mechanistic Personality Analysis of LLMs: Steering Personality via Latent Feature Interventions
- FacePlex: Toward Natural Full-Duplex Conversational Avatars
- Design and Implementation of Agentic Orchestrations and Orchestration of Agents
- OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3
- What is Left for Us? Second Scholarship Against the Degradation of Research by AI
- TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
- Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
- TopoExplore: Homology as an Exploration Signal
- Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Self-Improving Personal Agents
- The Steering Budget: Examples beat Knobs
- Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG
- SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
- DWM: Separating World Effects from Actions in Latent World Models
- Codifying the Judge: Scalable Evaluation via Program Distillation
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
- ClinLens: Towards Long-Horizon LLM Agents for Longitudinal Multimodal Clinical Data Science
- NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI
- MissClick: Execution-Aware Adversarial Attacks on Coordinate Generation in GUI Grounding Models
- Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs
- From Behavior to Mechanism: Tracing Divergent Response Modes in Frontier Language Models
- SDDBMs: Soft Denoising Diffusion Bridge Models
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- Causal Behavioral Evaluation of AI Agents at Scale via Automated Behavioral Science
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
- AeroCopilotBench: Safety-Gated Evaluation of LLM Agents on Aircraft Emergency Procedures in an Executable Cockpit
- What is Missing from AI Post-Training AI: An Empirical Analysis
- Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees
- Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents
- Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
- LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform
- From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents
- A Systematic Survey of Agentic Skills: Architecture, Lifecycle, and Security
- Escaping Reasoning Basin Collapse with History-Biased Search
- EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
- Fresh Memory, Stale Plans: Derivation Currency for Distributed LLM-Agent Memory
- DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
- SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
- What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
- Agents' Overreliance on Unreliable Tools
- A visual large language foundational model for medical image recognition using clinician-contributed online resources
- RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving
- CausalVerify: End-to-End Verification of Causal Analyses by Language Models
- A Theory of Reliable Self-Evolution for Agent Harnesses
- Structural Process Supervision for Latent Chain-of-Thought Reasoning
- A Dominant Supplier Slows Recursive Drift More Than It Steers It
- Do Not Restart: Residual Completion for Stateful Agent Handoffs
- DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents
- Enabling Creative Exploration for Vibe Design Agents
- Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election
- Reason What Matters: Retrieval-Grounded Reasoning for Universal Multimodal Embeddings
- AlgoEvo: Self-Evolving Agentic Search for Automated Algorithm Discovery
- The Pain Axis: LLMs Represent Self-Directed Harm and Act on It
- AquiLLM: Evaluating Faithfulness in Open-Weight RAG-LLM Systems for Scientific Research
- Compiled Agency: Coding Agents as Game AI Researchers -- from a Roguelike to StarCraft II and Civilization
- When2Think: Learning When and How Much to Reason
- DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
- What Should We Ask Next? Retrieval-Aware Question Learning for Interactive ReID
- Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
- Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA
- Generative Embodied Multiple Behavior Control Systems for Human-like Agents
- PINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks
- APEXA: Execution-Integrity Enforcement for Multi-Agent LLM Automation of Synchrotron Data Reduction
- The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
- Et Tu, Brute? Economic Misalignment in Personal AI Agents
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- Canonical locks that encode part-whole hierarchies
- JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
- CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
- EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
- Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution
- WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
- An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
- TRACER: Trajectory-Aligned Learning for Multi-Turn User Simulation
- Learned Cross-Task Relationships in Multi-Task Models
- Harness Tokenomics: A Router for the Enterprise Agentic Control Plane
- CounterRoute: Self-Routed Reasoning via Hierarchical Counterfactual Credit Assignment
- Breaking the Environment Wall: A Unified Framework for Preparing and Evolving Agent-Native Environments
- HEXIS: Compiling Agent Skills into Extended Finite State Machines
- Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
- UQ-LOB: Uncertainty-Aware Limit Order Book Mid-Price Forecasting
- Improving Causal Effect Estimation of Weighted RegressionBased Estimator using Neural Networks
- Training-Free Uncertainty Estimation for Embedding Models
- ELiSe: Efficient Learning of Sequences in Structured Recurrent Networks
- Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
- Detection and Characterization of Coordinated Online Behavior: A Survey
- Help Me Help You: The Aggregate Value of Source and Target Data in Transfer Learning
- Goal-Conditioned Supervised Learning for Multi-Objective Recommendation
- Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-based Robots
- Localized time-frequency representation learning for bioacoustic classification in complex soundscapes
- Learning Where and What to Restore for Composite Image Restoration
- Emergence of psychopathological computations in large language models
- Focus on Likely Classes for Test-Time Prediction
- Tags for DAGs: Graph Refinement with Meta-Informed Relations
- Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
- LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
- Bridging MOOCs, Smart Teaching, and AI-Assisted Learning: A Unified Quantitative Model
- On the Interaction of Compressibility and Adversarial Robustness
- RooseBERT: A New Deal For Political Language Modelling
- Graph Structure Learning with Temporal Graph Information Bottleneck for Inductive Representation Learning
- Learning Domain- and Class-Disentangled Prototypes for Domain-Generalized EEG Emotion Recognition
- One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
- Pushing Toward the Simplex Vertices: A Simple Remedy for Code Collapse in Smoothed Vector Quantization
- Dynamic Buffers: Cost-Efficient Planning for Tabletop Rearrangement with Stacking
- Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
- Graph Your Own Prompt
- Quantifying How Training Gradient Sparsity Affect Spiking Neural Network Accuracy And Robustness
- AttentionViG: Cross-Attention-Based Dynamic Neighbor Aggregation in Vision GNNs
- Learning to Generate Rigid Body Interactions with Video Diffusion Models
- Investigating The Smells of LLM Generated Code
- AI Adoption Across Mission-Driven Organizations
- NOSA: Native and Offloadable Sparse Attention
- Scaling Vision Transformers for Functional MRI with Flat Maps
- AlignBeat: A Latent Variable Model for Multi-Class Beat Tracking from Partially Labeled Data
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- Reinforcement Learning and Consumption-Savings Behavior
- Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems
- A Survey on Efficient Vision-Language-Action Models
- Sparse, self-organizing ensembles of local kernels detect rare statistical anomalies
- COGNOS: Universal Enhancement for Time Series Anomaly Detection via Constrained Gaussian-Noise Optimization and Smoothing
- D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Frequency and Pixel Spaces
- Server-Enforced Watermarking in U-Shaped Split Federated Learning
- MASTEST: A LLM-Based Multi-Agent System For Testing RESTful APIs
- The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs
- Why They Disagree: Decoding Differences in Opinions about AI Risk
- Time Series Foundation Models for Process Model Forecasting
- Soft Geometric Inductive Bias for Object Centric Dynamics
- FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
- SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
- Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search
- Deep Delta Learning
- NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
- Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs
- Contextual Distributionally Robust Optimization with Causal and Continuous Structure
- A Scalable Entity-Based Framework for Auditing Bias in Large Language Models
- OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
- Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
- MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
- Hypersolid: Emergent Vision Representations via Short-Range Repulsion
- Denoising Time Matters:Diverse Generation in Diffusion Language Models
- Plain Transformers are Surprisingly Powerful Link Predictors
- VLM-Guided Experience Replay
- MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
- Poly-attention: a general scheme for higher-order self-attention
- All-Atom GPCR-Ligand Dynamics Simulation via a Residual Latent Flow Mode
- Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
- Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
- Preventing Rank Collapse in Federated Low-Rank Adaptation with Client Heterogeneity
- ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
- CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
- Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
- Don't stop me now: How Validation Criteria Affect Checkpoint Selection and Early Stopping
- LFPO: Likelihood-Free Policy Optimization for Masked Diffusion Models
- Physics-Informed Neural Networks with Architectural Physics Embedding for Large-Scale Wave Field Reconstruction
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- A theoretical model of dynamical grammatical gender shifting based on set-valued set function
- SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models
- The Trace Is the State: Exact Credit Assignment for LLM Agent Teams
- Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
- RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
- Adaptive Weighted h-Transform Sampling for Coarse-Guided Visual Generation
- HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free Aggregation
- FlashSampling: Fast and Memory-Efficient Exact Sampling
- IndexRAG: Index-Time Reasoning for Multi-Hop Retrieval-Augmented Generation
- CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion
- Discovering What You Can Control: Interventional Boundary Discovery for Reinforcement Learning
- When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
- Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF
- P^2O: Joint Policy and Prompt Optimization
- Selective Deficits in LLM Mental Self-Modeling in a Behavior-Based Test of Theory of Mind
- The Model Says Walk: Measuring whether LLMs Condition on Hidden Constraints
- Meta-TTL: Meta-Learning Self-Improvement Policies for Language Agents
- AromaGen: Interactive Generation of Rich Olfactory Experiences with Multimodal Language Models
- An Inspectable LLM Council for Multi-Model Research Answer Aggregation
- How Does "English (US)" Become the Default? Triangulating Structural Bias Towards American English Across the LLM Pipeline
- Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment
- EffiPair: Improving the Efficiency of LLM-generated Code with Differential Execution Feedback
- An Imperfect Verifier is Good Enough: Learning with Noisy Rewards
- HiFloat4 Format for Language Model Pre-training on Ascend NPUs
- FastGrasp: Learning-based Whole-Body Control Method for Fast Dexterous Grasping with Mobile Manipulators
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- UniRect-CoT: Enhancing Generation in Unified Multimodal Models via Reflective Rectification with Inherent Understanding
- MambaSL: Exploring Single-Layer Mamba for Time Series Classification
- Why Fine-Tuning Encourages Hallucinations and How to Fix It
- Demystifying the Unreasonable Effectiveness of Greedy Alignment Methods
- Learning Evidence Highlighting for Frozen LLMs
- SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation
- DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- Query-Dependent Use of Generated Descriptions for Reliable Visual Question Answering
- Federated Semi-Supervised Graph Neural Networks with Prototype-Guided Pseudo-Labeling for Privacy-Preserving Gestational Diabetes Mellitus Prediction
- RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
- Deco: Extending Cherished Physical Objects into AI Companion Agents through Dual Embodiment
- MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
- Demystifying Manifold Constraints in LLM Pre-training
- Evolving Idea Graphs with Learnable Edits-and-Commits for Multi-Agent Scientific Ideation
- Two-Stage Learned Decomposition for Scalable Routing on Multigraphs
- One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent
- SymDrift: One-Shot Generative Modeling under Symmetries
- Pushing the accuracy of on-top functionals with agent-driven supervised learning
- When to Trust Imagination: Adaptive Action Execution for World Action Models
- Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
- UniPool: Learning Expert-to-Layer Ownership from Brief Global Access
- Enabling Unsupervised Training of Deep EEG Denoisers With Intelligent Partitioning
- Skip What You Can Predict: Predictive Repositioning for Policy Optimization for Efficient LLM Training
- On Privacy in Data-Space Tabular Diffusion Models: Influential Factors, Attacker Knowledge, and Metrics
- MELD: Multi-Task Equilibrated Learning Detector for AI-Generated Text
- In-Context Credit Assignment via the Core
- BGM-IV: AI-Powered Bayesian Generative Modeling for Instrumental Variable Regression with High-Dimensional Covariates
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Encoder-Decoder Transformers: Logical Characterizations and Periodicity
- Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
- GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
- One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy
- Tool Calling is Linearly Readable and Steerable in Language Models
- LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment
- AIPO: Learning to Reason from Active Interaction
- Recovering Physical Dynamics from Discrete Observations via Intrinsic Differential Consistency
- FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
- Agent Collectives Should Not Detect Their Own Imposters: A Chess Case Study
- Towards Effective Theory of LLMs: A Representation Learning Approach
- Neural Cluster First, Route Second: Capacitated Vehicle Routing via Differentiable Optimal Transport
- PoDAR: Power-Decoupled Audio Representation for Generative Modeling
- Investigating Single-Block Recurrence in Vision Transformers for Image Recognition
- When and How to Canonize: A Generalization Perspective
- Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
- Deep Minds and Shallow Probes
- Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
- Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems
- EHR-RAGp: Prototype-Guided Retrieval of Longitudinal Electronic Health Records for Clinical Prediction Models
- CoRe-Gen: Robust Spectrum-to-Structure Generation under Imperfect Fingerprint Conditions
- Watermarking Should Be Treated as a Monitoring Primitive
- Cross Modality Image Translation In Medical Imaging Using Generative Frameworks
- Where Should Diffusion Enter a Language Model? Geometry-Guided Hidden-State Replacement
- COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion
- EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
- Who Owns This Agent? Tracing AI Agents Back to Their Owners
- Goal-Conditioned Supervised Learning for LLM Fine-Tuning
- Reducing Hallucination in Multimodal Large Language Models through Hard Grounding Preference Supervision
- The End of Trust: How Agentic AI Breaks Security Assumptions
- DeepArrhythmia: Segment-Contextualized ECG Arrhythmia Classification via Selective Evidence Acquisition
- EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
- Fidelity Probes for Specification--Code Alignment
- ContractBench: Can LLM Agents Preserve Observation Contracts?
- DCFold: Efficient Protein Structure Generation with Single Forward Pass
- Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models
- OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
- When Does Equivariance Help? Canonical Alignment in Neural Fluid Surrogates
- MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
- Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
- DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Agent Memory
- Anytime Training with Schedule-Free Spectral Optimization
- XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
- One-Forcing: Towards Stable One-Step Autoregressive Video Generation
- PhoneWorld: From Real-App Trajectories to Dynamic and Verifiable Environments for Phone-Use Agents
- EviLink: Multi-Path Schema Linking with Uncertainty-Guided Evidence Acquisition for Large-Scale Text-to-SQL
- Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents
- No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
- Shared Doubt: Zero-Shot Cross-Lingual Confidence Estimation for Language Models
- BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization
- Can Predicted Dynamics Exist in the Physical World?
- Context-aware tokenization for Cross-subject Emotion Decoding from EEG
- Knowledge-Intensive Video Generation
- Quantifying and Mitigating Domain Shift in Peach Leaf Damage Classification: Attention Mechanisms and Fine-Tuning Strategies
- ConTraIRL: Factorized Contrastive Abstractions for Transferable IRL
- From 'What' to 'How' and 'Why': Sharing LLM-Generated Retrospective Summaries of Older Adults' Passive Tracking Data with Remote Family Members
- Channel-Preserving Representation Alignment for EEG-to-Music Reconstruction
- Large Language Models Hack Rewards, and Society
- Dual Advantage Fields
- Vault: One-Step Latent Generation with Positive-Anchored Rewards for Autonomous Driving
- Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks
- Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
- Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems
- Agentic Search for Counterfactual Recourse under Fixed LLM Budgets
- Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis
- When Good Verifiers Go Bad: Silent Negative Transfer in Verifier-Guided VLM Training
- Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning
- Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces
- AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents
- StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
- Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition
- What Should a Streaming Video Model Remember?
- When Does Depth Matter For In-Context Learning? Adaptive Inference in Deep Transformers
- DRIFT: Data Selection for LLM Instruction Tuning via On-Policy Attribution
- Guava: Distilling Frontier VLM Agents into a Compact Model with a Manipulation Harness
- When to Commit and When to Defer: Maturing Markov Decision Processes under Refining Information and Expiring Opportunities
- OneCanvas: 3D Scene Understanding via Panoramic Reprojection
- CRAX: Fast Safe Reinforcement Learning Benchmarking
- Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG
- Rethinking Object-Centric Representations for Video Dynamics Modeling
- JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
- A Gravitational Interpretation of Safety Reversion under Fine-Tuning
- Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
- Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
- Scaling Laws for Collapse in Asynchronous GRPO
- Sampling Meets Interaction: Sequentially-Controlled Multi-Particle Flow-Maps for Efficient Inference-time Search
- LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
- BFMT: Enhancing Search Capabilities of Tree Sampler via Bootstrap Flow-Map Tree
- One Image, No Tokens: A Controlled Study of Glyph-Based Chinese Language Modeling
- Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
- Source-Lifted Flow Matching for Intervenable Multimodal Imitation
- Towards Autonomous and Auditable Medical Imaging Model Development
- Agent Hacks Agents: Autoresearch Discovers Vulnerabilities in Production Agents
- The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- What Does Model Growth Really Add? Functional Capacity Beyond Parameter Count
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM Pruning
- CriPO: Enhancing Rubric-based RL via Self-Distillation
- MOPDA: Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection
- Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
- Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs
- Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction
- Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity
- DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
- One Analyst Is Not Ground Truth: Grading Agent-Built Financial Models Against Observed Professional Practice
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
- SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
- RedFlow: Redirect Failure into Action-level Corrections for Flow-matching VLA Policy
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
- DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction
- Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
- dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
- Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
- E$^3$-Orch: Towards Effective, Efficient, and Extensible Agentic Orchestration with Reinforcement Learning
- SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
- Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
- Multivariate Time Series Forecasting needs Cross Variable Loss
- FastKron: Efficient Quantization with Kronecker-Factored Hessians
- TransSLR: A Lightweight Transformer for Sign Language Recognition
- MaskFlow: Precise, Consistent and Seamless Regional Image Editing
- Evidence-RL: Towards Evidence-intensive Visual Reasoning
- RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
- RefineFly: Failure-Aware Post-Training for Aerial Vision-Language Navigation
- Error-Aware Reverse Auction Mechanism for Large Language Model Routing
- ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
- AQuA: Recursively Self-Improving Quantitative Trading Research Agents
- Spectral Saliency for Machine Unlearning
- Multinomial Subset Routing with Sum-Max Rewards and Operational Constraints
- PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX
- Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
- In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography
- Luce: Relightable Gaussians for 3D Asset Generation
- Fusing Perceptual Vision Experts with Multimodal LLMs for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation
- FRAME: Separating sampling variation from performance disparities in medical image analysis
- Flow-JEPA: Robust Latent Dynamics for JEPA World Models via Flow Matching
- History-Conditioned Joint-Prefix Alignment for Generative Recommendation
- PokaiTrainer: Scaling Equilibrium Search to Competitive Pok\'emon VGC
- Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals
- AGM: Achievement-Grounded Memory for Closed-Loop Agents with Frozen VLA Policies
- EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction
- HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives
- Linear Reusable Neural Bases Architecture for Network Compression
- Percolation Dynamics in Optimization: Variance Cascades and Nested Symmetry
- Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
- Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
- Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
- TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables
- Beyond the Matrix Sign: Quadratic Spectral Descent
- KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
- In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Post-Retrieval Context Tampering
- Break Step: Recursive Training Resonates with Replayed Sampling Noise
- Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer
- Rice's Theorem under Self-Modification: Elevation Operators and a Normal Form
- SkillAtlas: An Attack Trace Library for Agent Skills
- A latent dimension of Condorcet's jury theorem for multiple AI advisers
- WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians
- Neural-Network Solutions to Real-Space Charge Density and Generalization
- SpliTEE: Fast and Private LLM Inference by Coupling GPU-Assisted Trusted Execution Environments with Differential Privacy
- Refinement-Based Flow Policy Optimization
- Evaluating Losslessness in Speculative Decoding Under Finite-Precision Inference
- K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
- The Router Within: Eliciting Native Skill Routing from a Frozen LLM
- Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents
- The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
- RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- Toward Personalized Sleep Guidance from Wearable Data Using Language Models
- Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
- Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?
- Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements
- Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification
- Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
- The Undetected Damage of Quantization on Retrieval and How to Fix It
- iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
- Teaching a Moving Student: Rethinking the Curriculum of On-Policy Distillation
- Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes
- What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
- In-Context Guidance: Learning Inter-Task Synergies via Numerical Foundational Models for Few-Shot Multitask Optimization
- QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models
- Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning
- Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
- GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning
- Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations
- MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
- Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations
- Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions
- Softmax Reparameterization for Output-Head Quantization
- UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
- Replay in the Silent Degrees of Freedom: Continual Learning Without an Offline Phase
- OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit
- EEGAgentBench: Benchmarking LLM Agents on Short- and Long-Horizon EEG Analysis
- Enhancing generalization in endwall film cooling prediction: Incorporating the superposition principle into transformer-based neural operators
- Symmetry-quotient Flatness and Generalization
- Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning
- FIDAL: Diversity-Aware Federated Active Learning Under Real-World Distribution Shifts
- Product-Aware Deterministic Rounding for Quantized Matrix Multiplication
- Cross-Material Support Transfer for Core-Loss Prediction Under Waveform Covariate Shift
- Beyond the Graph: An Adaptive Meta-Learner Fuses Explainability, Weather, and Dynamics for Robust Bus ETA Prediction
- NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization
- When Keywords Drop but Classifiers Hold: Soft Refusals under KV Cache Compression
- 3-D Emissions Mapping and Social Cost Estimation for US Domestic Aviation at West Coast Hubs
- seq2cause: One Autoregressive Backbone, Four Causal Discovery Tasks in Event Sequences
- Medium-Term Multi-Resolution Electric Load Forecasting using Economic Data and Foundation Model
- Relational Compression: A Framework for Relational Fidelity in Constrained Representations
- Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problem
- TemporalGraphLLM: Temporal Graph Neural Networks with Large Language Models for Dynamic Text-Attributed Graphs
- Resource-Aware Federated Mixture-of-Experts with Adaptive Pruning for Onboard Learning in LEO Satellite Constellations
- Cache-Aware Conv3D Lowering Across Embedded World-Model Decoders
- Simple Extensions of Single-Objective Acquisition Functions and Hedge Strategies for Multi-Objective Bayesian Optimization
- On-Policy Attention Linearization
- Model-Agnostic Online Certificate-Driven Calibration for Time Series Forecasting Under Distribution Shift
- Model Casting and Low-Parameter Gating: Towards More Sparsely Activated FFNs
- Understanding the Subspace Stabilization of the Hessian and Gradient Covariance Matrix
- Lagrangian and Hamiltonian Neural Networks With a Dissipative System
- Mechanistic Interpretability Reveals Shared Causal Subspaces in Brain-to-Speech Decoders
- Can Circuit Alignment Predict OOD Generalization?
- Integrating Language Models into Listened and Imagined Speech Decoding from MEG
- Human Activity Recognition via Ultra-Wideband Data: A Framework for Dimensionality Reduction, Pattern Discovery, and Predictive Modeling
- Representation Learning for Exact Preimages
- Graph Forward Distribution Matching for Molecular Inverse Design
- Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox
- LLM Unlearning Evaluation with TRIAGE
- An Attention-Driven Heterogeneous GNN Model for Credit Card Fraud Detection
- SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models
- How Reusable Are Benchmarks with Richer Feedback?
- Handwritten Digit Leakage from Smartphone Motion Sensors Across Unseen Users and Phone Models
- What Should We Freeze? Guarded Freezing: Connectivity Shapes the Fine-Tuning of Pretrained Models
- Latency-Aware Client Assignment for Parallel Split Learning With Global Sampling
- Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
- CAFE: Counterfactual Prediction via Fast Posterior Estimation
- Efficient Support Recovery of Mixtures of Sparse Linear Classifiers with Less Measurements
- Analytic-Walk Rotary Positional Encodings for Graphs
- Rethinking Cross-Channel Importance in Time-Series Forecasting
- HM-ROUTER: Joint Model and Harness Routing for Agentic Systems
- CompassPlay: Rewarding the Proposer for Where It Moves the Solver
- "Where Can I Trust You?": Boundary-Aware Evaluation of Surrogate Fidelity
- Arithmetic Simplicity in Stochastic Gradient Methods
- Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes
- Representation Editing for Multimodal Test-Time Adaptation
- When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift
- Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
- Certification Frontiers for Gaussian LoRA: Independent Priors, Posterior Risk, and Prediction-Preserving Balancing
- Learning Through Game: Skewed Transfer of Tabular Knowledge to Strengthen Image Model
- Topology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor Index
- HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
- SIMANF: Sample Free Learning of Unnormalized Distributions via Simulated Annealing in Normalizing Flows
- A Journey to the Edge of Stability
- FUND: Density Flow for Sampling Unnormalised Distributions
- Refresh or Realize? Compute Allocation in Drifting Models
- When Does Synergy Help Active Feature Acquisition? A PID-Based Study
- On the Capability and Limitation of Hard Prompt
- Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
- Continual Data Unlearning in Diffusion Models via Transition-based Regularization
- When the Merge Coefficient Stops Mattering: Proximity Regularized Merging for Continual LoRA Adaptation
- A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
- Editable Map-Conditioned Trajectory Generation for Human Mobility Simulation
- Graph Memory: Spectral Associative Memory via Dirichlet Energy
- STRIDE: State-Transition Representation via Increment Dynamics and Evolution
- Attribution Without a Second Pass: Inline Per-Sample Gradient Provenance at ~1% Overhead
- Using Machine Learning to Investigate Predictors of Fasting Blood Glucose: Insights into Circadian Timing and Age Interactions
- Recovery-Directed Symbolic Distillation of Neural Likelihoods
- HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration
- Beyond the Manifold Hypothesis: Hybrid Spectral Parameterizations for Flow Matching
- Length-Independent State Tracking Under a Parallel Scan
- Write Back the $\Delta$: Revisiting the Same Tokens with Fresh Representations
- PolyStepOR: Learning to Decide Without Optimal Decisions
- On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
- Fast and Precise Learned Charged-Particle Trajectory Regression at the Large Hadron Collider
- What Should Federated LoRA Share? FedSAIL via Input-aware Subspace Alignment
- Elastic Selective Spectral Hybrids for Train-Once, Export-Many Budgeted Inference
- Neural Dynamics as the Composition of Quantized Units
- Fisher Simplicity in Kolmogorov-Arnold Networks and Multilayer Perceptrons
- What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response
- Does Transolver really need a Transformer?
- Shared Autoregressive Context Can Distort Relationships in Synthetic Data
- Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery
- Compositional Objectives: Learning Structure in Structure
- CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization
- Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy Distillation
- DimPO: Dimensionality Reduction for Attention using Preference Optimization
- Age of Learning: Temporal Persistence of Prediction Errors as a Learning Signal
- Learning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node Classification
- Stabilizing the Dynamic Low-Rank Training
- Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching
- Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution
- Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
- Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
- SIFT: Enhancing Time Series Foundation Models via Semantic Invariance and Structural Fidelity Fine-Tuning
- Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning
- BiasReducer: Adaptive Bias Mitigation for Reward Models
- Scaling Properties of Same-Family On-Policy Distillation
- Distributionally Robust Average-Reward Reinforcement Learning: Finite-Sample Guarantees under Weak Communication
- Predicting the Financial Impact of Supply Chain Risk for Major AI-Related Semiconductor Firms: A Heterogeneous Graph Patch Transformer Approach
- Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations
- Reuse or Relearn? A Spectral View of Earth Observation Foundation Models
- AnchorMixGAN: Anchor-Aligned Generative Semi-Supervision for DDoS Detection in Cloud-Integrated IoT Networks
- Structuring Relations Among Learning Paradigms via Protocol--Objective--Resource Reductions
- Beyond Gaussian Assumptions: Distribution-Aware Channel Capacity for Effective Connectivity
- Robust Bayesian Optimization with Q-Exponential Surrogates
- Learning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental Learning
- FedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective Optimization
- Understanding and Exploiting Anisotropy in Post-Training
- Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
- Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
- Retimed Bellman Flows: Escaping the Impossible Triangle of Velocity Bootstrapping
- UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
- Over-the-Air Federated Learning in Heterogeneous Mobile Wireless Networks
- Permutation-Equivariant Flow Matching for Alignment-Free Neural Weight Generation
- HamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic Maps
- Muon Under Gradient Noise and the Limits of Orthogonalization Near Optima
- Staying on the Attractor: Supervising Neural Surrogates of 3D Turbulence Where They Leave It
- Measurement-Gated Provenance Attenuation for Frozen EEG Representations
- A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks
- Predicting the Next State Is Not Enough: JEPA Representations for Lean Theorem Proving
- Adaptive Latent Capacity for World Models
- Phenomenon-Graph JEPA: Label-Efficient Representation Learning for Contactless Cardiorespiratory Sensing
- Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation
- Counting on Thinking: Tracing Evidence Integration in Language Models
- Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs
- The Impact of Stochasticity on the Rashomon Effect in Machine Learning
- Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination
- Generative Priors Conditioned on Natural Language for Bayesian Inversion in PDEs
- Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
- Constrained Flow Policy Updates: A Generalized Schr\"odinger Bridge View
- Clipped or Unclipped? Finite-Sample Trade-offs for Averaged SGD under Heavy-Tailed Noise
- Self-Confirming Superposition Traps in Reinforcement Learning
- Adaptive Ensemble Selection for Noisy Labels on Tabular Data
- Feasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual Guidance
- What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
- Saturation-Insensitive Dueling Bandits with General Function Approximation
- Low-Rank Single-Index Bandits with Unknown Links: From Matrices to Tensors
- What Must a World Model Distinguish for Planning?
- Balancing Early Performance Sacrifices with Long-Term Gains: Scaling Learning-Rate Warmup Duration Across Training Horizons
- Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers
- DevelopmentODE: Structured Neural ODEs for Early Brain Development Dynamics Across a Decade
- SketchSSM: Write to the Full State, Read from a Compact Sketch
- Deep Learning Techniques for Phoneme Recognition in Italian Children' s Speech
- KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
- The limits of exactness: On the failure of automatic differentiation in physics-informed machine learning
- From Constitutions to Control: Interpretable Rewards for Aligning Language Models
- Geometry-Aware Operator Families for Structured Representation Learning
- CARVE: Breaking Data Barriers in Chip Placement by Harnessing Reusable Expertise
- Simulation-Free Learning of GP-SDEs from Irregular Observations
- Mycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution Systems
- Save Your Saturated Data: Learning Beyond Reward Saturation in Group-Based RL
- ILP-BO: Integer Linear Programming-Based Black-Box Optimization
- Leaky Students: Membership Inference against On-Policy Distillation
- CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation
- Convergence of Practical Muon
- Downstream-Aware Context Selection for Online In-Context Reinforcement Learning
- When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability
- Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems
- When an Evaluation Rule Writes Training Labels: Measuring Human-Reference Forgiveness in NAVSIM
- AG-CoT: Verified Algorithmic Traces for LLM Program Synthesis on Clifford Circuits
- Apparent Compression, Real Stability: The Intrinsic Dimension of Learning a Quantum Wavefunction
- Orthogonal Witness Control for Muon Optimization via Sigmoid Spectral Reshaping
- Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
- RMB: Reward Model Boosting Mitigates Reward Hacking
- MTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring Systems
- The Price of Locality: Why Forward-Forward Underperforms Backpropagation?
- Minimax-Optimality of Posterior Sampling for Reinforcement Learning
- Feedback-Robust AI for Patient Knowledge Graphs
- Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts
- GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences
- Towards Identifiable Representations under Misspecified Structure
- Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
- When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression
- Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning
- RAEGL: Risk-Aware Evidence-Gated Learning for Selective Contextual Routing under Temporal Shift
- CalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property Prediction
- MultiEcho: An Experimental Science of Learned Worlds
- How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure
- The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection
- TNF based Spectral Embedding for Effective Application of Supervised Machine Learning Techniques in Automobile Insurance Fraud Detection
- Optimal Transport Dropout for Structured Predictive Uncertainty
- From Grey-Box to Green-Box: When can Physics-Informed Machine Learning Reduce Carbon Footprints in Structural Health Monitoring?
- StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks
- FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward
- MoGround: Measuring and Mitigating Modality Distraction in Vision-Language Models
- SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval
- SMAT: Simple and Efficient Merge-Aware Training
- Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning
- A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation
- Geometric Identification in Predict-Then-Optimize Learning
- How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage
- What masking geometry works best for EEG foundation models?
- Predicting Block-Coordinate Performance via Cross-Curvature
- Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
- The cost of useful natural gradient updates
- Pulseflow: PPG Counterfactual Generation Via Latent Transport
- SLP-ProbHard: Probabilistic Hard-Constrained Learning via Structural Latent Parameterization
- LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation
- Discovering Symmetries in Neural Network Parameter Spaces
- Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving
- Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile
- GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference
- Correct then Forecast: Observer State-Space Models for Time Series Forecasting
- HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
- Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration
- TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
- Pretraining Transformers with Quantized Softmax in Attention
- Short-Length Code Designs for Integrated Sensing and Communications: A Deep Learning Approach
- Hierarchical Response Preservation for Continual Adaptation of Zero-Shot Graph-Text Models
- You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
- Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
- Quantifying Behavioral Tails in Black-Box Language Models
- SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models
- FuseAlign: Forced Alignment in the Wild
- Towards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning Tasks
- Scalable Attribution and Control of Model Behavior During Training
- Reachability is not enough: Diagnosing long-range behavior in GNNs
- Benign Overfitting for General Norms and Distributions
- T-MoXAI: A Hierarchical Explainability Framework for Temporal Multimodal Data
- Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer
- A Spectral Theory of Compositional Learning
- DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation
- Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning
- ALDER: Discovering the Laws of a World by Acting in It
- PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention
- Task-Aware Discretization of Differentiable Logic Gate Networks
- ResDiffFRG: Residual Diffusion for Multiple Appropriate Facial Reaction Generation
- Oracle-Efficient Online Classification with Stochastic Inputs and Adversarial Outputs
- StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences
- Binding Multiple Modalities via Multimodal Wasserstein Barycenter
- dOPT: Differentiating Conic Optimization via Geometric Reduction
- ChemOPD: Multi-Teacher On-Policy Distillation for Multi-Task Chemical Reasoning
- QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
- Rethinking Contextualization by Reinterpreting Attention Head Channels
- JET: Justification Evaluation in Transformer
- Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits
- Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
- Augmented Feature Boosting for Multicalibration
- MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
- No Free Efficiency: Revisiting the Trade-off Between Training Efficiency and Model Vulnerability
- On the Two Faces of Adam in Separable Linear Classification
- Training Witnesses: Trusting the Training without Trusting the Trainer
- Optimizing the Phi-2 Small Language Model for Real-time Chatbot Applications Using Parameter-Efficient Fine-Tuning (PEFT) with QLoRA Quantization
- How Strong Is the Evidence for the Artificial Hivemind? Reevaluating Evidence for the Open-Ended Homogeneity of Language Models
- Behavioral Monitoring of JEPA World Models with Jacobian Centroids
- DynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly Detection
- Future Information-Directed Sampling for Bayesian Nonstationary Bandits
- From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting
- ICMAPE: In-Context Multiagent Pure Exploration
- ASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite Constellations
- T-SNN: Temporal Simplicial Neural Network for EEG Decoding
- RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting
- Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
- LTV-CTDNet: Compositional Turning Decomposition for Short-Term Turning-Movement Forecasting
- Structure-Adaptive Tree Field Integrators
- Posterior Regimes and Latent Deception: Variational Bayesian Inference in Hidden Markov Models for Sequential Fraud Detection in Financial Transactions
- Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
- PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
- Steering Language Model Goals with Value Transplant
- KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
- Counterexamples to Local Reconstruction Gain as a Proxy for Final Fidelity in Residual Completion
- Beyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation Models
- When Is an SAE Feature Interpretable? A Validation Ladder for EEG Foundation Models
- Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
- Evolution of fairness in multi-objective reinforcement learning framework
- What Does a Stream Model Buy You in Flow Matching?
- Training and Inference Dynamics of PLDR-LLMs: Row-Map Collapse, Renormalization, and Predictive Reduction
- ExpertoRhythm: Morphology-Aware Learning for Waveform Reconstruction and Cuffless Blood Pressure Estimation from Single-Channel PPG
- SPINET: Sheaf Protein Inverse Folding Network
- Transfer Calibrated Prediction Powered Inference
- Hidden Activations are not Enough I: Knowledge Matrices as Higher Representations
- GPARA: Graph-Posterior-Aligned Refinement and Active Acquisition for Grounding Diffusion Priors
- Cardinality-Stratified Interaction Decomposition for Interpretable Pairwise and Higher-Order Structure in Transactional Basket Data
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
- Learning to Optimize through Solver-Grounded Self-Play
- X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths
- Loop Dropout: Regularizing Shared Updates in Looped Language Models
- SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals
- CasEm: A Cascade Architecture for Long-Horizon Neural Emulation
- DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models
- Rotated Manifold Optimization for Low-Rank Adaptation
- ZeroCode: On-demand Error-Correcting Code Construction from the Zero Matrix via Reinforcement Learning
- Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
- Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search
- Epistemic Learning from Imprecise Annotation
- Riemannian Difference-of-Convex Optimization for K-Means Clustering
- One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
- Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions
- The Composition Gap in Dataset Distillation
- FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model
- FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents
- Spexis: Speculative Lookahead Scheduling for LLM Inference
- Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
- GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
- Unlocking Few-Step Diffusion for Faithful Previews
- PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
- MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism
- On the Relation Between Interval Regret and Dynamic Regret
- Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
- LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
- Harmonizing Spectral Evolution in Conditional Flow Matching for TTS
- Admissible Diffusion for Multimodal Interventional Trajectories
- Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths
- Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
- ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models
- PhysioTRACE: Provenance-Aware Stress Tests for Physiological Foundation Models
- Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning
- Deep kernel hedging
- When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
- M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
- LLN: Learnable Lens Networks for Parameter-Efficient Long-Horizon Dynamical Prediction
- QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
- Distribution-Conditioned Task Routing for Class-Incremental Learning
- KiT: A Foundation Model for Financial Time-Series Forecasting using DiffusionTransformers
- Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs
- HALO: Enhancing Time Series Generation via Hyperspherical Latents and Masked AutoregRessive Modeling
- On Parameter Symmetries and Conservation Laws in Gradient Flow
- Verifying Neural Networks with Reinforcement Learning
- PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding
- Brain-Conditioned Action Policies for Neural Motor Decoding
- Single-Layer MeMo as a Randomized Hamming-Kernel Classifier
- Retracing Hodgkin and Huxley: State Recovery Does Not Certify Mechanism
- Compute Time Scaling with Recursive Models for Combinatorial Optimization
- Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation
- The Low-Rank Structure of VLA Reinforcement Learning
- When local gains fail to transfer: Frozen Earth-observation embeddings across wildfires
- Shaping Persistent Representations from Independent Interactions
- PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
- Beyond Site Agreement: Re-estimation for Brain Network Generalization
- Learning Regional Snow Water Equivalent and Snow Height Variations from Sentinel-1 InSAR Acquisitions
- Evaluating Dynamical Fidelity through Predictive Structure in Physical Representations
- DisKO: Deep Koopman Learning in Distribution Space from Unpaired Snapshots
- GenMem: Generative Symbolic Memory for Self-Evolving Harness
- Correction-space Cross-variate Interaction for Test-time Adaptation in Time Series Forecasting
- Uniform Race: Parameter-Free Approximate Rejection Sampling
- Tilted Schr\"odinger Bridge Matching
- Universal Dynamic Portfolios
- When Can Attention Heads Be Statically Defined?
- Minimax Last-Iterate Convergence in Matrix Games with Observed Actions
- FestDPO: Few-step Generator Alignment with Direct Preference Optimization
- QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
- SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
- AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
- Predicting Delayed Train Trajectories on the Dutch Railway Network: Explainable AI Evaluation of Topological, Operational and Weather Features with Tree Based Ensemble Methods
- Learning Propagation Geometry from Message-Passing Feedback
- Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning
- Edge-Level Automorphism in GNNs: A Quantitative Framework and Effective Designs For Link Prediction
- Instance-Adaptive Prompts as Context for Time-Series Foundation Models
- Separating personal from population gains when calibrating EEG foundation models for new users
- Polylogarithmic Nash Regret in Matrix Games with Bandit Feedback
- Structured Neural SDEs for Functional Calibration
- QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
- Context-dependent time-series prediction via HyperReservoirs
- When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
- Learning High-Risk High-Precision Motion Control
- Muon Sublates the Edge of Stability in LLM Pretraining
- Sample What You Say: Aligning Language Models to Sample the Distributions They State
- Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models
- XMatch: Enhancing Covariate-Aware Time Series Forecasting through Tree-Structured Exogenous Matching
- Adjoint Guidance Flow: Amortized Critic Guidance for VLA Policies
- ALICE: In-context, Zero-shot, Mutual Information Estimation
- See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs
- Teach to Learn: Hint Annealing for Self-improving LLM Reasoning
- Inspector: Conversational and Lightweight Analyzer of Analog Circuit Layouts Using LLM and CNNs
- ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
- MaPP: A Unified Marginalized Posterior-Predictive Framework for Data-Efficient RLVR
- Price Stability in the European Union: A Systemic Approach Using Random Matrix Theory
- Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons
- THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout
- Universality and Generalization of Causal Transformers Across Context Lengths
- Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
- Depot-Closed Multi-Component Construction for Neural Vehicle Routing
- Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning
- Interrelating Fruchterman-Reingold Graph Visualization and Agglomerative Clustering
- ReCo: When to Relocate Sensor Kits under Deployment Constraints -- A NILM Case Study
- Propagate, Then Sharpen: Post-Hoc Refinement of Frozen Node Classifiers
- Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
- GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents
- Retrieval-Augmented Diffusion Modeling for Stochastic Discount Factor Portfolios
- E3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPU
- DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
- SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
- A Multimodal Autonomic Sensing Framework for Objective Assessment of Patient Responses to Dental Pulp Stimulation
- Explaining Hyperbolic Neural Networks via Geometry-Aware Relevance Propagation
- FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
- CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion
- Small transformers track Bayesian evidence for latent common causes via a context-invariant mechanism
- EdgeCraft: Automated Model Crafting for Edge IoT
- ProtoSeam: Lifting Classifier Training with Latent Gaussian Mixture Models
- Subgroup Rank-1 Lattice for Practical High-dimensional Black-box Integral Approximation
- ConRAG: Lightweight inference of multi-hop relations
- Adversarial Consistency-Guided Representation Learning for Multi-view Clustering
- Temporal Heterogeneous Graph Pretraining for Relational Deep Learning
- TANGO: Watermarking Masked Diffusion Language Models in Token Pairs
- Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
- Disentangling Lung-Cancer CT/LDCT AI: A Systematic Evidence Map of Clinical Tasks, Evidence Chains, and Translational Gaps
- AIM-ZO: Activation-Informed Subspace Maintenance for Zeroth-Order LLM Fine-Tuning
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
- Latency and accuracy tradeoffs in Spiking Neural Networks
- SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards
- $\lambda$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
- Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
- LionMuon: Alternating Spectral and Sign Descent for Efficient Training
- When Should a Satellite Estimate Be Changed? Stress-Testing Neural Corrections for Evapotranspiration
- Collaborative Principle Evolution via Evidence Transfer for Scientific Discovery
- Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data
- Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
- NeuronDiscover: Agent-in-Twin for Mechanistic Discovery in Neuronal Microenvironments with World Action Models
- Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
- Quasi Linear Kernel Attention with Infinite Capacity
- Interference Beyond Geometry in Concept Extraction
- Fiona: Accelerating FHE Inference with Packing-Aware Ternary Weights
- First Learn, Then Memorize: The Spectral Bias of Diffusion Models
- Identifying Neural Source Dynamics from Unknown Local Interventions
- Inductive Feedback for Mixed-Policy Distillation
- Persistent Partners Raise Prices Among Learning Agents
- LLMs are General Asynchronous Agents
- SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
- NeuronSifter: Intervention Planning in CNS Microenvironments
- Manifold-Stable Flow Matching
- Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
- An analysis of Mirror-Descent Soft Actor-Critic
- Universal Approximation of Measure-to-Measure Operators by Pushforwards
- Physics-Guided Conditional Diffusion Model for Rare Event Synthesis and Diagnosis for the Water-Gas Shift Reaction
- Structured Latent Modeling for Supervised Multimodal Information Decomposition
- An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
- One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
- Reward-Aligned Reweighting for On-Policy Distillation
- Deep Epistemic Value Functions for Optimistic Exploration
- GeoGAE: Scalable Graph-Level Autoencoding via Hyperball Cloud Representations
- Optimal Networks for Agentic Information Aggregation
- Learning the Robustness Mechanism with Bilevel Optimization
- Simplex Diffusion Models
- From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis
- Output-aware Residual Stream Pruning for Large Language Models
- Hardware-Aware Features for CUTLASS Kernel Selection
- Control-Geometry Straightening for Sampling-Based Latent Planning
- EvE: An Alternate Optimizer to Adam
- Attention Graphons: A Graph Limit Perspective on Graph Transformers
- Cartridges++: KV Cache Compression without Off-Context Derailment
- Arbitrary-Accuracy Neural Approximation with Optimal Neuron Count and Near-Optimal Bit Complexity
- SANTA++: Sampling Attention through Representative Keys
- Bounding Retraining Equivalence and the Deletion Floor in Materials Machine Unlearning
- Transferable Mass Spectrum Prediction via Reference-Guided Test-time Specialization
- Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models
- Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
- MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
- ScAn-Bench: Evaluating Scaling Analysis Methodology
- Neural Harmonic Measure Operator
- Unifying Distributional Training for One-Step Visual Generation
- Kolmogorov-Arnold networks in nuclear binding energy prediction
- Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations
- Exterior complex scaling enables physics-informed neural networks for quantum scattering
- FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes
- Accurate Sampling from Diffusion Models
- Distributional sentiment modeling and anomaly detection for consumer complaint assessment
- Learning Steadily: Accumulating Relative Point Margin Scores for Face Image Quality Assessment
- The Temporal Tug-of-War: Visualizing and Detecting RAG Conflicts in Diffusion Models via Trajectory Variance
- Data Processing for Offline Evaluation in Recommender Systems: a Survey
- Be Careful Who You Trust: Coordination Dynamics under Corrupted Communication in LLM Multi-Agent Games
- Statistical Testing for Multiple Instance Learning via Selective Inference with Applications to Computational Pathology
- Seeing the Heat: Synthesizing High-Resolution Wood Thermal Responses from Optical Imagery
- Calibration-Free Surface Normals Estimation in Vision-Based Tactile Sensing using Universal Photometric Stereo
- Acoustic domain shift in spoken language identification from systematic domain generalization evaluation to real-world application
- Suitable Measures for the Potential Operational Utility of AI NWP Rainfall Forecasts Over Africa
- Panoptic Scene Program Diffusion Transformer
- Cross-Modal Knowledge Distillation for Acoustic Pedestrian Detection
- Convergence-Aware Pareto Selection of Covariate Scaling Transformations for Markov Deterioration Hazard Models: Evidence from Bridge Inspection Data
- Prompting Particle Physics: Tokenized Multi-modal Foundation Models for Combinatorially Many Tasks
- Knowledge-Driven XRD Phase Identification via Multi-View Retrieval and Explanation
- MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus
- GT-VLA: Target-Conditioned Trace Guidance for Generalizable Robotic Manipulation
- Recipe-Matching, Not Equivalence
- Artificial Neural Network Assisted Modelling of Tangent Galvanometer Measurements for the Determination of Horizontal Component of Earth's Magnetic Field
- Transformer MLP Gate Thresholds Are Couplings to a Carried Reference Direction
- Bridging Stochastic Flow Maps and Boltzmann Generators with Normalizing Flows
- Evasion Attacks on Cost-Utility-Based Adversarial Training for Online AutoML in IoT Networks
- Identifiability Limits of Gravitational Wave Phase Deviations: Multiclass Classification with a Multihead Neural Network
- Continuous-Time Trajectory Generation from Discrete Observations with Stochasticity
- Correcting the Dropout-LayerNorm Expectation Gap Improves Protein Structure Models
- Period Segmentation in Transition Network Analysis: A Topological Data Analysis Approach
- A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets
- FOCUS: Fixed-Confidence Online Causal Learning Using Sequential Adaptive Interventions
- Uncertainty Quantification of Next Generation Reservoir Computing with Applications to Memory-Driven Dynamical Systems
- Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models
- High-Probability Guarantees for SGD under $\beta$-Heavy-Tailed Gradient Noise
- The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing
- SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs
- Kernel-Based Steering of CLIP with Vision-Language Model Preferences
- DegreeSpar: Structured Degree Sparsity for Efficient Secure Transformer Inference
- Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation
- Active Data Acquisition with Side Information via Discrete Diffusion Priors
- Before Answering: Evidence Sufficiency under Size-Matched Memory Construction
- Overfitting of Spectral Gradient Descent: How Matrix Geometry shapes Generalization and Implicit Bias
- Neural ODEs Meet Concurrent Learning: Stable Online Learning with Lyapunov Guarantees
- All On-Board: Fully On-Chip Neuromorphic Q-Learning with Embedded CartPole Simulation
- Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
- Gradients for Interventions and Activations for Detection: Targeted Feature Learning in Language Models
- Convergent Plug-and-Play Image Restoration with Annealed Noise Levels
- AECSF: Adaptive Ensemble Conditional Score Filtering for High-Dimensional Nonlinear Data Assimilation
- Masking Frequent Tokens Sharpens Direct Preference Optimization
- What Should the Reflector See? An Empirical Study of Evidence in Reflective Prompt Optimization
- Agnostic Smoothed Online Regression with Adversarial Responses
- JEPA Learns What the Mask Leaves Unrecoverable
- When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections
- Fast Differentiable SVD on GPU via Polar Decomposition
- Language as an Independent Information Layer: A Conceptual Model of Communication, Cognition and Decision-Making
- Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
- HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion
- Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC
- KV-Lingo: Learning KV-Cache Translators with Distillation
- Learning What to Evaluate: Correlation-Aware Decoupling for Multiobjective Bayesian Optimization
- Schur-Neural KF: Learned Schur-Consistent Corrections to the Extended Kalman Filter
- Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection
- Affine Geometry of Gaussian ReLU Networks via Conditional Kac-Rice Formulas
- One-Step Generative Modeling via Unbalanced Optimal Transport
- Revisiting AdaGrad in Stochastic Convex Optimization: Last Iterates, High Probability, and Lower Bounds
- Domain Adaptation with Target Information via Doubly-Anchored Distributionally Robust Optimization
- Plan-to-Synthesis: Cross-City Human Mobility Generation via Semantic Latent Flow Matching
- SAGE: Semantic Audio Generative Encoder
- An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning
- Copper-Policy: Focus on the Representation for Robust Robot Manipulation
- OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing
- CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models
- CollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion
- OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
- USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments
- Learning Shuffle Ideals with Membership Queries and Contrastive Examples
- Mend the Measurement Gap: Latent User Preference Modeling for Short-Form Video Recommendation
- VCRE-Fib: View-Conditioned Regional Evidence for Fine-Grained Ultrasound Grading of Schistosoma japonicum-Associated Liver Fibrosis
- SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals
- Sliced Orlicz-Wasserstein
- Is H&E Image-to-Spatial Transcriptomics Simpler Than It Looks?
- Machine learning for the LHC physics program: a 2025-2026 stocktake
- Neural Network-Assisted Refinement of Traditional Schemes for One-Dimensional Scalar Conservation Laws
- EEG-Based Motor Imagery BCI Algorithms and Technologies: A Review
- The Key Handoff: Retrieval in Hybrid Language Models
- Distributed Hydrological Modeling in the Feature Space
- Decentralized Optimization with Cross-Coupled Mixed Affine Constraints
- Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction
- Reading Too Much into Context: Passive Exposure Can Steer LLM Decisions
- Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality
- Evolving Dexterous Robots from Scratch
- Can Tabular Foundation Models Amortize Statistical Inference?
- CAME: Company-Aware Evidence-Memory Experts for Interpretable Quarter-Ahead Revenue Forecasting
- Sharp Critical Minimax Laws and No-Learning Thresholds in Continuous-Time Adaptive Control
- Notes on Generative Modeling for Feedback Control and Planning
- Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall
- Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations
- MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers
- Schr\"odinger--F\"ollmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample Analysis
- Modelling non-linear aeroelastic loads in long-span bridges with extreme learning machines
- The Statistical Benefits of Multiple Responses for Learning from Demonstrations
- When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions
- Geometry-Adaptive Mechanisms for Private Synthetic Data
- Identifying the Predictable Drift of a Semimartingale from Marginal Laws
- Local LMO is Secretly a Projection Method!
- Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors
- Recovering Lower-Dimensional Semialgebraic Support of a Measure from its Moments
- A Visual Classification Dataset and Model Evaluation for Historical Manuscript Illustrations
- Sharp training-conditional coverage for conformal prediction under covariate shift
- Mask-Induced Displacement in Audio XAI via Logit Trajectory Decomposition
- Domain-Adapted Diffusion Models for Conditional Independence Testing
- Neural Scaling Laws of Transformer Operator Network
- Calibrated Derivative-Process Sensitivity for Gaussian-Process Variable Selection
- Sparsity by Default: The Theory and Practice of ARD in Gaussian Process Regression for Variable Selection
- Terminal-Register Certification for Finite-Measurement Learning of Multiscale Quantum States
- Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
- The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces
- lapanda: A Matrix-Free Differentiable Solver for Nonconvex Constrained Optimization Layers
- LieDiscover: Adaptive Symbolic Library Construction for Explicit Open-form Symmetry Discovery
- COGNIT-Guard: Calibrated Standalone Direct-Decision Guardrails with Heterogeneous CPU-NPU Confidence Cascading under Explicit Latency and False-Positive Constraints
- Prompt-Anchored Residual Adaptation for Biomedical Vision-Language Models
- Non-Adaptive Learning of Sparse Erd\H{o}s--R\'enyi Graphs via Affine Splitting
- Reliable Replay through Spatial Coherence in Online Continual Learning
- A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions
- Constrained Edit Fields for Training-Free Flow Editing
- Weighted Spline-Expanded Networks with Distributional Balancing for Continuous Treatment Effects
- Multi-Marginal Inverse Optimal Transport for Contrastive Learning Via Explicit Anchor-Positive-Negative Coupling
- YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
- Positions Are Not Facts: The Mismatch Between KV Caches and Memory
- EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
- Taming the Greeks: Option Portfolios with Inductive Biases
- Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
- RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation
- Autonomous phase discovery
- Controlling Speaking Rate in Autoregressive TTS via Activation Steering
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
- Annealed Sinkhorn with Momentum: Certified Unregularized Optimal Transport in Linear Memory
- One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
- CLIMB-flow: Coupled Linear Inverse posterior sampling via Multiscale-Based flow
- ViBR-WM: Visual Bayesian Regression for World Modeling
- DCEmbed: Scalable Optimization over Neural Surrogates
- Neuron-Level Architecture Growth: A Controlled Evaluation for EEG Time-Series Decoding
- Prospective Interpretation Risk: Principled Communication Control Between LLMs
- Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
- Residual-Stream Burden Shapes Representation Learning in Diffusion Transformers
- Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
- Greenpixie's AI Token Methodology: Assessing the Energy, Water and $\mathrm{CO_2\text{-}eq}$ Impact of AI Tokens for Open and Closed Weight Models
- Two-Sample Testing for Inhomogeneous Random Graphs in Non-Integral $L_r$ Norms
- Adapting neural operators for mechanics decisions under changing operating conditions
- High-Level Text Preprocessing for Semantic Similarity Analysis of Discursive Texts: A Framework and Empirical Demonstration
- The Privacy Fallacy of Crowdsourced Fine-Tuning: Extracting Proprietary Data via Topic-Based Poisoning
- RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
- Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control
- 3D Point Tracking with State Space Models
- Singularities of Non-negative Matrix Factorization and their application to Bayesian inference
- Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
- The Statistical Cost of Causal Discovery with Feedback
- The Double-Edged Sword of Information: Revealed versus Hidden Lotteries in School Choice
- Forecast-Necessary Causal Discovery for Nonlinear Political Panel Data: Feedback, Functional Form, and the Dynamics of Democratization
- Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems
- GT-PSSM: Unified Probabilistic Framework for Stochastic Dynamics Modeling and Dependency Learning in Multivariate Time Series Anomaly Detection
- Functional Autoencoders for Amplitude-Phase Representation Learning
- Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
- RoboICL: Embodied In-Context Learning with GPT-6 Astra
- Dexterous Tactile World Model
- Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale
- Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
- Unbiased Top-$k$ Estimation for On-Policy Distillation
- ARS: Agentic Reward System for Robot Learning
- How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models
- RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
- AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum
- The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
- Probabilistic Geodesic Flow Matching on Location-Scale Families
- CLAD: Constrained Abstract Domain for Neural Network Verification
- On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation
- Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks
- Hierarchical Clustering and Signal Denoising on Digraphs
- Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
- Information-Theoretic Analysis of Next-Token Prediction under Markovian Data
- Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
- Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks
- Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models
- Finite-Time Concentration and Convergence Rates for Projected Two-Time-Scale Stochastic Approximation with Markov Noise
- Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
- On Temporal Binding in Large Audio Language Models
- From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models
- Accelerator Choice Is Not Enough: AlphaFold2 Inference on Cloud TPUs
- MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation
- When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model
- Conformal Prediction and Conditional Coverage for Tabular Foundation Models
- Physics-Informed Neural Networks for Depth-Averaged Avalanche Dynamics
- N\"urnberg NLP at ChildSafeAds 2026: Structurally Dissimilar Voter Ensembles under Four Levels of Data Access
- Simulation-Based Quantum System Inference with Neural Posterior Estimation
- Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
- Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition
- GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization
- Graph-Based Learning for Multi-Horizon Martian Atmospheric Forecasting
- Perceptual Quality Loss or Loss of Perceptual Quality?
- Continuous Variational Synthesis
- DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection
- Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach
- G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA
- Uncertainty Quantification in Cardiac Model Personalisation from Ultrafast Ultrasound
- A Hierarchy of Entropy-Shapley Games for Multivariate Predictive Uncertainty
- Beyond Selection: Token Parameterization for Extreme Visual Token Compression
- AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
- Scaffold Then Internalize: Representation Injection for Diffusion Transformers
- Simulation-Based Inference for Plate Reverb System Identification
- Convex Optimization Is Free When Accuracy Is Expensive
- Multi-Task Learning of Conditional Mean Operators: applications to dynamical systems and uncertainty quantification
- From internal representations to model improvement through prediction errors
- Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
- TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training
- Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
- Learning Conditional Expectation Operators via Functional Newton Updates
- On-Policy Self-Distillation for Multi-Turn Image Editing
- Elicitation and Decision Geometry in Single-Index Bandits
- Which the Eye Fears: Writing with Read-Blindness Explains Massive Activations in Transformers
- CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
- Learned Preconditioning for a Primal-Dual Interior-Point Method
- The Hidden Perception Constraint in Task-Aware Compression
- Harness Learning Enables Generalizable Test-Time Adaptation
- Improving Test-Time Scaling with Adaptive Looped Transformers
- Statistical Learning of Contractive Dynamical Representations for Composite Adaptive Control
- PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
- Advanced Policies: A First-Principles Path from Policy Gradient to Q-Learning
- FedSysID: A Federated Approach to Sample-Efficient System Identification
- Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
- Classifier-pruned Bayesian optimization for particle accelerator tuning: Exploring temporally structured manifold of 6D beam phase space
- Stochastic Engrams for Efficient Continual Learning
- DRAN: A Distribution and Relation Adaptive Network for Spatio-temporal Forecasting
- AYLA: Architecting a loss landscape in shallow neural networks to accelerate feature recovery
- LOD: Latent Objective Discovery in Heterogeneous Multi-Agent Reinforcement Learning
- Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization
- Understanding and Mitigating Under-Confidence in GNNs from the Final Layer
- On the $O(\frac{\sqrt{d}}{K^{1/4}})$ Convergence Rate of AdamW Measured by $\ell_1$ Norm
- Relation-Aware Graph Foundation Model
- Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning
- HERO: Preserving Structure and Semantics in Heterogeneous Continual Graph Learning
- Federated Independent Component Analysis via Spectral Alignment and Robust Aggregation
- AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
- Edge Selection for the Effective use of Piecewise-Constant Distributions as Neural Network Outputs for Event Prediction
- SHEFL: Sparse Heterogeneous Ensemble Federated Learning with Group-Balanced Aggregation
- An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
- WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning
- How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
- Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models
- Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion
- Federated Computation of ROC and PR Curves
- Auditing Information Disclosure During Large-Scale Gradient-Based Training via Gradient Uniqueness
- Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
- Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
- SCORENF: Score-based Normalizing Flows for Sampling Unnormalized distributions
- Managing Self-Learning Experts under Per-Round Budget Constraints
- A Game-Theoretic Spatio-Temporal Reinforcement Learning Framework for Collaborative Public Resource Allocation
- Wavelet-Based Parity Detection Revisited: Representation Dependence, Generalization, and Mechanistic Analysis
- Neural Tractability via Structure: Learning-Augmented Algorithms for Graph Combinatorial Optimization
- CAT: Can Trust be Predicted with Context-Awareness in Dynamic Heterogeneous Networks?
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- Role Support in Knowledge-Graph Error Ranking: Predictor Regimes and Evaluation Policies
- Quantifying the Effect of Test Set Contamination on Generative Evaluations
- Understanding and inverse design of implicit bias in stochastic learning: a geometric perspective
- Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning
- Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
- DP-{\lambda}CGD: Efficient Noise Correlation for Differentially Private Model Training
- Environment-Conditioned Tail Reweighting for Invariant Learning under Mixed Shifts
- Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment
- $f$-FUM: Federated Unlearning via min--max and $f$-divergence
- SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
- PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
- Watch the Model Think: On-Policy Extraction of Activation Steering Vectors
- Pseudo-differential-enhanced physics-informed neural networks
- Use What You Know: Causal Foundation Models with Partial Graphs
- Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generative Modeling
- Forecasting Bacterial Antimicrobial Resistance Trends Using Machine Learning on WHO GLASS Surveillance Data: A Retrieval-Augmented Generation Approach for Policy Decision Support
- Models Designed to Forget: Machine Unlearning via Key Deletion
- Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
- Binary Classification from Coupled Pairwise Labels
- Joint Surrogate Learning of Objectives, Constraints, and Sensitivities for Efficient Multi-objective Optimization of Neural Dynamical Systems
- Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics
- LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
- Reasoning Shift: How Context Silently Shortens LLM Reasoning
- Contrast Matters: Understanding Robustness of In-Context Fine-Tuning to Target-Context Relatedness
- Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
- Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
- EvoLen: Evolution-Guided Tokenization for DNA Language Model
- Model Compression with Exact Budget Constraints via Riemannian Manifolds
- A Theory of Saddle Escape in Deep Nonlinear Networks
- Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
- Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies
- AeroJEPA: Learning Semantic Latent Representations for Scalable 3D Aerodynamic Field Modeling
- From Dual Tracking to Clipping: Provably Faster Distributionally Robust Multi-Objective Optimization
- Learning beyond Site Bias for OOD Generalization in Brain Networks
- LINC: Decoupling Local Consequence Scoring from Hidden Matching in Constructive Neural Routing
- Diverse Sampling in Diffusion Models with Divergence-Free Particle Guidance
- Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse
- Why Does Agentic Safety Fail to Generalize Across Tasks?
- CellScientist: From Execution Feedback to Auditable Model-Revision Trajectories for Cellular Perturbation Prediction
- On the Invariance and Generality of Neural Scaling Laws
- Self-Play Enhancement via Advantage-Weighted Refinement in Online Federated LLM Fine-Tuning
- Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
- FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning
- Let the Target Select for Itself: Data Selection via Target-Aligned Paths
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
- DynaMiCS: Fine-tuning LLMs with Performance Constraints using Dynamic Mixtures
- ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
- A Dual Representation of Influence Functions for Linearizable Models
- Hypernetworks for Dynamic Feature Selection
- PyroAdapt: Adapting Wildfire Prediction under Spatial Heterogeneity and Temporal Shift
- Hessian Matching for Machine-Learned Coarse-Grained Molecular Dynamics
- Learning to Select Source Domains: Proxy-Rewarded Policy Optimization for Molecular OOD Generalization
- Gaussian Relational Graph Transformer
- To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Learning, Except In Heavy Truncation Scenarios
- AMO: Operator-level Adaptive Muon Orthogonalization
- Beyond Square Roots: A Memory-Efficient Explicit Factorization for Multi-Epoch Private Learning
- PACE-FNO: Physics-Aligned Canonical Equivariance for Fourier Neural Operators
- Learning Robust Recommenders from Noisy Implicit Feedback via GMM-Weighted Bayesian Transition Matrix
- Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
- EntmaxKV: Support-Aware Decoding for Entmax Attention
- Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
- Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
- Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models
- From Privacy to Generalization: Linear Max-Information Bounds for Differentially Private Learning Algorithms
- Inference-Native Zeroth-Order Optimization for LLMs
- Digitally enriching a high-risk population for pancreatic cancer using routine blood-based measures and clinical histories
- Effective Biological Representation Learning by Masking Gene Expression
- SPR: Toward a Graph Foundation Model for Transferable Graph Cognition via Spectral Patterns and Relational Geometry
- The Routing Plateau: Understanding the Accuracy Limits of LLM Routers
- nCMD: Benign-Anchored Feature Selection for Imbalanced Network Intrusion Detection
- AugRelNet: Relation-Augmented Dynamics for Structure Discovery and Forecasting from Limited Data
- Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
- Multi-Bitwidth Quantization for LLMs Using Additive Codebooks
- Running the Gauntlet: Hard Agentic Tasks
- RepNN: Tackling spectral bias in deep neural networks for regression and PDE problems via parameter reparameterization
- Filtered Conformal Ellipsoids for Graph-Native Time Series
- Recursive Scaling in Masked Diffusion Models
- CODEBLOCK: Learning to Supervise Code at the Right Granularity
- When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness
- Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment
- The Interplay of Harness Design and Post-Training in LLM Agents
- Fast LeWorldModel
- EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix
- The Weakest Link Tells It All: Outcome-Supervised Process Reward Modeling via Learnable Credit Assignment
- Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias
- Revisiting the Volume Hypothesis
- What Does a Routing Oracle Measure Under Stochastic Decoding? Coupling, Scorer Choice, and Single-Commit Ceilings
- Transformers with Physics-Informed Encodings and Simulation-Based Inference for Robust Detection of Eccentric Binary Black Holes in Pulsar Timing Array Data
- RL Forgets! Towards Continual Policy Optimization
- Higher-Order Geometric Updates for Levenberg-Marquardt Method via Riemann Normal Coordinates
- Hyperbolic Manifold Constrained Tabular Neural Network
- Extractable Memorization From First Principles
- A Continuous-Time Reinforcement Learning Framework for Fine-Tuning Discrete Diffusion Models
- Robust Peak-cost Constrained Reinforcement Learning
- Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting
- Conditioned Direct Feedback Alignment via Activity and Error Geometry
- AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
- Interval and fuzzy physics-augmented neural networks (iPANN and fPANN) for uncertainty quantification and propagation in constitutive modeling
- On the Convergence of Stochastic Low-Rank Adaptation
- Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram
- Variational Boosting for Physics-Informed Neural Networks
- Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality
- What Can Latent World Models Know? Physical Information in Multimodal Predictive Representations
- THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
- DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
- SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization
- Introspecting Alignment Shifts Beyond Behaviors Implanted Through Fine-Tuning
- Predicting Task Difficulty Without Rollouts
- A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies
- Which Decisions Low-Bit Quantization Breaks, and How to Predict Them
- Memory--Batch Tradeoffs in Lipschitz Bandits
- HOPPER: Learnable Hop Extraction for Linearized Graph Sequence Models
- From Objectives to What Models Learn: A Landau Theory of Invariant Learning
- Convergence Guarantees of Gradient Descent for Neural Networks via Generalized Lipschitz Smoothness
- FLARE++: Low-rank attention with dynamic attention routing
- Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
- A Reproducibility Study of Partial Residual Ablations in Pre-LN Transformers
- Online Convex Optimization with Dueling Feedback
- SAGA: Structure-Attended Generative Action Embedding Model that encodes Multi-Surface User Action Sequences
- SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces
- J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules
- Reinforced Planning with Latent World Models
- Beyond Multimodal Alignment: Shared Physical Representations Across Sensors and Action Orders
- Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra
- Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Neuromorphic Virtual Sensing
- Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
- Seasonality-Aware Hybrid Convolutional Transformer for Antarctic Sea Ice Concentration Forecasting
- Nonparametric Contextual Pricing and Inventory Learning under Censored Demand
- How Do Language Models Choose Between Context and Memory?
- TRACE: A spatiotemporal contact memory graph network simulator for granular dynamics
- Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate
- No-Regret Mixing of LRU and LFU with Optimal Switching Cost
- PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games
- Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
- JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
- Pathwise Individual Rationality in Federated Learning: A Mechanism-Architecture Co-Design
- Sharp Regret Bounds and a Task-Covariance Correction for Spectral Representation Learning
- Scaling Laws for Physics-Aware ACOPF Surrogate Learning
- Stiefel Attention: When the Geometry of Transformer Projection Matrices Dominates Optimizer Choice---and When It Does Not
- Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks
- Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining
- PRQuant: Permutation Residual Quantization for Low-Overhead Inference
- A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
- CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era
- Graph Domain Adaptation Does Not End with Representation Learning
- PR-Smoother: Simulator-Preserving Non-Gaussian Smoothing for Data Assimilation
- Physics and Data Driven Transformer-Mamba Framework for Flow Field
- Beyond Feature Reliability: Repeat-Informed Multifractal Curve Regression for Brain-Age Prediction
- FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates
- Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation
- RAZOR: Pruning Replaceable Experts in LLMs
- Benchmarking Attention for Tabular Foundation Models
- Bridging Generative and Discriminative Noisy-Label Learning via Direction-Agnostic EM Formulation
- Damped Gauss Newton Search for Multi Metric Hyperparameter Optimization
- Revisiting Inexact Fixed-Point Iterations for Min-Max Problems: Stochasticity and Structured Nonconvexity
- LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
- Generalization Analysis of Online Stochastic Gradient Descent for Overparameterized Two-Layer Neural Networks
- A Survey on Reinforcement Learning Applications in SLAM
- TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
- Coreset-Based Task Selection for Sample-Efficient Meta-Reinforcement Learning
- Optimization on the Oblique Manifold for Sparse Simplex Constraints via Multiplicative Updates
- Towards Weaker Variance Assumptions for Stochastic Optimization
- FEAT: Free energy Estimators with Adaptive Transport
- When do Random Forests work?
- Accelerating Natural Gradient Descent for PINNs with Randomized Numerical Linear Algebra
- Squeeze3D: Extreme Neural Compression with Latent Space Bridging
- A family of graph GOSPA metrics for graphs with different sizes
- Identifiable Convex-Concave Regression via Sub-gradient Regularised Least Squares
- Adaptively Truncated Signature-based Logistic Regression for Semi-parametric Functional Classification
- AI-driven Dispensing of Coral Reseeding Devices for Broad-scale Restoration of the Great Barrier Reef
- Design of Experiment for Discovering Directed Mixed Graph
- Synthetic data for ratemaking: imputation-based methods vs adversarial networks and autoencoders
- A Kernel-based Stochastic Approximation Framework for Nonlinear Operator Learning
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Systematic Exploration of Multi-core Architectures for Efficient LLM Serving using WaferAI-SIM
- Large Language Model Selection with Limited Annotations
- R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages
- Action-Driven Processes for Continuous-Time Control
- Disciplined Biconvex Programming
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Digital Agriculture Sandbox for Collaborative Research
- Reinforcement Learning to Initialize Newton-Raphson for AC Power Flow with Quantum Annealing-Based Environment Updates
- Manifold limit for the training of shallow graph convolutional neural networks
- Model-Free Output Feedback Stabilization via Policy Gradient Methods
- Instance-optimal high-precision shadow tomography with few-copy measurements: A metrological approach
- Efficient reduction of stellar contamination and noise in planetary transmission spectra using neural networks
- TADA! Tuning Audio Diffusion Models through Activation Steering
- Adjoint-based shape optimization of a ship hull using a Conditional Variational Autoencoder (CVAE) assisted propulsion surrogate model
- Symmetric Composition of Anisotropic Operators for Global Subseasonal-to-Seasonal Climate Forecasting
- Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
- SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution
- Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing
- Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
- When Should Humans Step In? Optimal Human Dispatching in AI-Assisted Decisions
- Patient-level validation of a foundation-model pipeline for pediatric lung sounds: physician-annotated adventitious events are recognized, disease groups are not reliably predicted
- Q-Drift: Quantization-Aware Drift Correction for Diffusion Model Sampling
- The Exponentially Weighted Signature
- Fast-Slow Thinking RM: Efficient Integration of Scalar and Generative Reward Models
- Hyperspectral Trajectory Image for Multi-Month Trajectory Anomaly Detection
- Stringological sequence prediction I: efficient algorithms for predicting highly repetitive sequences
- HRIR-Former: Grid-Free Time-Domain Reconstruction of Head-Related Impulse Responses with a Spatially Encoded Transformer
- Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- Conformal Robust Set Estimation
- Sliced-Regularized Optimal Transport
- Multiscale Euclidean Network Trajectories: Second-Moment Geometry, Attribution, and Change Points
- Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
- TrajGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations
- Probabilistic Object Detection with Conformal Prediction
- ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation
- Broximal Gradient Descent: A Projection-Free Sister of Projected Gradient Descent
- Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
- Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration
- FIS-DiT: Breaking the Few-Step Video Inference Barrier via Training-Free Frame Interleaved Sparsity
- Synthetic American Option Pricing via Jump-HMM-Driven Heston Implied Volatility
- Regret Equals Covariance: A Closed-Form Characterization for Stochastic Optimization
- CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage
- Generalized Functional ANOVA: A Complete Theoretical Framework
- Uncertainty-aware neural emulation reveals climate-resilient maize traits at scale
- Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
- Equivalent Flows, Unequal Learning: Clean-Latent Prediction in Transformers
- Beyond Imitation: Reflective On-Policy Self-Distillation for LLM Reasoning
- Diagnosing Visual Ignorance in Vision-Language Models
- Driving Video Retrieval for Complex Queries with Structured Grounding
- Robust Active Learning for Few-Shot Example Selection in Text-to-SQL
- Trainability of IQP Quantum Circuit Born Machines Under Gaussian Initialization
- Toward Proactive RF Charging Scheduling: Generative AI for Decision Support
- DroneShield-AI: A Multi-Modal Sensor Fusion Framework for Real-Time Autonomous Drone Threat Detection, Behavioral Intent Classification, and Swarm Intelligence in Contested Airspace
- Capability Provenance in Language Models: A Case Study in Social Reasoning
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Finite-Sample Performance of Gradient Descent in Logistic Regression with Gaussian Design
- Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
- Retrieval Observability Bounds on Provenance Detection for Agent Memory Poisoning: Measured Coverage and a Falsified Standalone Detector
- Dithered Gaussian Mechanism for Randomness-Efficient Differential Privacy
- Artificial Intelligence Across the Cardiac Amyloidosis Diagnostic and Management Pathway: From Single-Modality Detection to Multimodal Clinical Integration
- NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations
- Diversified Multinomial Logit Contextual Bandits
- BARS: Benign-Anchored Ranking and Selection for False Alarm Reduction in Network Intrusion Detection
- Priors learned from legacy reconstructions inherit undetectable overconfidence
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Fractional Parabolic Partial Differential Equations in Anisotropic Spectral Barron Spaces: Regularity and Neural Approximation
- TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking
- Tokenizer-Generator Coupling in Medical Image Generation
- WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing
- Neural Boltzmann Equations
- Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision
- Scaling Reinforcement Learning for Diffusion Models via Velocity Matching
- Affix Cache for Diffusion Large Language Models
- Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers
- WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading
- PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion
- ActionSplice: In-Flight Action Editing for Interactive World Models
- MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining
- Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead
- Tight Sampling Complexity with Stochastic Gradient Oracles in Fixed Dimensions
- Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling
- Decentralized Gossip Learning and Federated Averaging for Histopathology Image Classification
- VLA-ULAP: Interleaving Cloud VLA Calls with Ultra-Lightweight Local Action Prediction at the Edge
- Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation
- Near-Optimal Single-Loop Predictor--Corrector Extragradient Method for Strongly Convex--Strongly Concave Minimax Optimization
- Complete Neural Electronic Initialization Accelerates Materials DFT
- Learning and Control Beyond Linearity: Towards a Non-asymptotic Theory for Bilinear Systems
- RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers
- A Bayesian Vertical Federated Learning Framework for Multivariate Reduced-Rank High-Dimensional Regression
- PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
- FREESIA: Covariance-Aware Posterior Transport for Expressive and Scalable Data Assimilation
- GINIO: A Geometric SO(3)-Equivariant Interface for Neural Inertial Odometry
- SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation
- Statistical Gains from Looped Estimation under Parameter Budgets
- Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
- How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?
- Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3
- HClimRep-Ocean: A Global Ocean Emulator on an Unstructured Mesh
- Learning a Flow to Self-Supervised Representations
- Cost-Sensitive Online Window Size Selection for Portfolio Management
- Low-Rank Friction for Memory-Efficient Transformer Pretraining
- Hypothesis-awkward: Property-Based Testing Strategies for Awkward Array
- Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories
- CG-Diff: Organizing Code Changes Around Call Graphs
- IChart2Code: Benchmarking Multimodal Large Language Models for Interactive Chart Code Generation
- How AI Changes DevOps Performance: A Mechanism-Based Simulation
- Beyond the Model: Demystifying Harness Effects in Software Engineering Agents
- Evolution of Deprecated APIs and Their Replacements in Python Libraries
- Social3Source: Connecting Social Provenance, Societal Qualities, and Social Purpose on an Open-Source Software Foundation
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
- Path2Spec: Path-Aware Specification Generation via Large Language Models
- Neuro-Symbolic Indirect-Call Analysis under Opaque Pointers
- Beyond the Prompt: Linking What Developers Ask, Do, and Understand with Coding Agents
- Green AI: Cost of LLM-Based Code Completion
- The Construction of an Empirical Dataset of Incomplete Software Changes from Open Source Projects
- The Last Mile Is the File: OfficeEditBench for Preservation-Aware Office Editing
- From Noisy Telemetry to Actionable Warnings: GPU Failure Prediction in Industrial Clusters
- When Ambiguity Meets Atypicality: Dual-Perspective Test Input Prioritization for DNNs
- Multi-SWT-Bench: A Multilingual Benchmark for Reproduction Test Generation
- TraceLib: System-Call Bitmap Feedback Mechanism for Language-Agnostic Web Fuzzing
- Certified Compilation in the TELEPERM XS Nuclear Safety I&C Platform
- Verifying Graceful Degradation in a Distributed Malware-Detection System with SPIN
- SmartMemory: Detecting On-chain-off-chain Communication Inconsistency for Smart Contract via Memory-based Agent
- Making the Invisible Visible: A Framework for Reflective AI Use in Software Engineering Education
- Is there a future for models in the LLM era?
- Distributed Service Orchestration in Edge-Cloud Continuum for Digital Healthcare
- Verification of PETSc with CIVL using LLM-generated ACSL contracts and deterministic driver generation
- PDFa11yMut: Measuring Mutation-Specific Detection in PDF Accessibility Checkers
- NxM-Version Programming for Quantum Software: High-Level Components across Frameworks and Engines
- EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning
- InfoEdit: Probing Global Layout Reasoning in Infographic Editing
- Faultless: A Program Equivalence Technique for Validating and Evaluating Neural Decompilers
- Semi-automated Verification of Symbolic Invariants In Extended Symmetric Nets
- Trajectory-Level Security Debt in LLM Coding Agents
- LLM-Assisted Automatic Security Proofs for Cryptographic Protocols: How Far Are We?
- The Cost of AI-Assisted Coding: An Extensive Investigation of Energy vs. Accuracy in Language Models
- SynthCoder: Anti-pattern identification and model training for FIM mode code completion
- Faster but Not Wiser: GitHub Copilot Decouples Programming Performance from Code Comprehension in Brownfield Tasks
- Beyond Accuracy: Behavioral Dynamics of Agentic Multi-Hunk Repair
- Large Language Models for Unit Test Generation: Achievements, Challenges, and Opportunities
- Clarity Is Not Assumed: Understanding LLM-Based Code Generation under Ambiguous Requirements
- Software Product Line Engineering: Adoption, Tooling and AI Era Challenges
- Do Newer Models Produce Better Patches? A Non-Functional Quality Study
- SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work
- TasmScan: Continuation-Aware Taint Analysis for TVM Bytecode with Savelist Abstraction
- A Set-Theoretic Evaluation Framework for Assessing Asset Administration Shell Instances: Towards Comparability and Suitability
- A Carbon-Aware Quantum Computing Framework for LCA-Driven Sustainability in Quantum Cloud Services
- Trajectory-Aware Benchmark Subset Selection for Cost-Efficient Software Engineering Agent Regression Testing
- Source-Known Identifiers: A Three-Tier Identity System for Distributed Applications
- Evolution of Log-Based Detection Rules in Public Repositories
- From Evidence to Effect: Authority Semantics and Runtime Infrastructure for Stateful Agents
- CoW — a stacking window manager for Wayland
- Pining for Arc Downcasting in Rust
- Aho-Corasick Algorithm
- Bill Gates tries to install Movie Maker
- My experience writing automated tests for a SPA
- What Would A Serious AI Product Look Like?
- Switching To Emacs as a Neovim User
- Yes, no AI is now a feature
- AI Didn’t Make Programming Easier. It Just Made It Differently Difficult
- MotifCentral - All Things Motif/Xt/Xlib
- Getting root on OnePlus 15 from an untrusted app
- Hijacking the PS5's RTMP Stream
- Leaving them behind
- State of the (Tagged) Union Address by Andrew Kelley
- When did Google get so weird?
- How I Built an iPhone App in Four Days with Opus 5.5
- Deterministic Concurrency
- What makes Lisp difficult to read?
- Text-to-meowdio models
- It’s Time to Investigate the AI Labs
- HardenedBSD August / September 2026 Status Report
- Rickrolling with a Pharmacy cross
- Adding Floating-Point Decimals for Fun and Profit
- Unix File and Directory Permissions and Modes (2018)
- Prototyping a Small Genetic Algorithms Library inHaskell (2019)
- Why I Replaced Vector Search with Hindsight for Meeting Notes (いいね相当スコア: 0)
- Exposing Fake Coin Receipts: A Public Security Breakdown | Ahmad Zorph (いいね相当スコア: 0)
- What Four Sessions of Live SAP Demonstrations Taught Us (いいね相当スコア: 0)
- Enforcing Zero Network Egress with Automated Sovereignty Tests (いいね相当スコア: 0)
- AI-Assisted Debugging Techniques for Complex Systems (2026) (いいね相当スコア: 0)
- Still Looking for My First Developer Job in the Age of AI 🚀 (いいね相当スコア: 1)
- Day 2 (いいね相当スコア: 0)
- Building a Deal Intelligence Agent with Persistent Memory (いいね相当スコア: 0)
- What changed when our cameras started remembering, with Hindsight (いいね相当スコア: 0)
- Listen to your coding agent instead of reading it: how I save my eyes (いいね相当スコア: 0)
- How I Stopped Bioreactor Batch Losses Using Hindsight Memory (いいね相当スコア: 0)
- How I Gave a Support Agent Memory With Hindsight (いいね相当スコア: 0)
- One Model, Many Roles: Specializing LLM Inference Without Training More Models (いいね相当スコア: 0)
- Making Agent Memory Visible (いいね相当スコア: 0)
- "Deal Whisperer" — Enterprise AI Sales Intelligence with Hindsight Vector Memory & Groq (いいね相当スコア: 0)
- Reverify weighs verified AI claims by how informative they are (いいね相当スコア: 0)
- Built a local AI pipeline that captures system audio, transcribes it, structures the result with a local LLM, and sends reviewed notes to Obsidian. I also wrote about what broke when real users started testing it. (いいね相当スコア: 0)
- Gita AI (いいね相当スコア: 0)
- Sonnet 5.5 did this. Opus 5.5 quality with half price.
- Opus 5.5 vs Sonnet 5.5 : 3D steampunk whale modeling
- I use a Claude pro account for work, my employer now wants me to share the login for that account with multiple coworkers. Will the account be banned if I do this?
- Are we living in the good old days of AI?
- Goat Riding a Tractor Compare Models - turns out that level of effort Max seems to be a deciding factor (mostly)
- With Opus 5.5 and Sonnet 5.5 both apparently outperforming Sol and Astra, Anthropic has technically made OpenAI’s Dev Day a lot more interesting. OpenAI is reportedly planning 20+ launches tomorrow, so I’m really curious to see what they have in store now. The timing couldn’t be more interesting. 😅
- We finally start praising Opus 5.5 and Claude goes down
- Claude Code can message easily between sessions now. This is a very powerful feature
- With Fable 5.1 being my main driver I kept hitting 80% of total usage on my $200 plan a week. With Opus 5.5 its about 25%. Have other people noticed the improvement?
- Anthropic offering discounts to old users — anyone else receive this? (IPO push?)
- I gave opus 5.5 one prompt about AI fear-mongering. It made this entire music video in code: research, lyrics, animation, a dancing 3D hologram, sound design. I didn't write a line.
- Opus 5.5 - is the Superpowers skill still needed?
- Its beautiful
- Anthropic response to Instinct / Muse / OpenAI Dots?
- I open-sourced the cozy pixel-art game Opus 5.5 made for me & my friends: here's the repo and a playable demo
- I got tired of pgAdmin, so I built my own database GUI
- Sonnet 5.5 is great and all but where to use it?
- Got my most expensive coding AI Sub i ever gotten after trying Alternatives
- Sonnet 5.5's prompt suggestion box literally says "Nothing obvious to suggest."
- From Superpowers to Superbrainstorming
- Claude Code desperately needs a read-only mode [Product Request]
- Weekly Self Promotion Thread
- Coding agents write a new helper instead of finding the one you already have
- how to code with astra without hitting limits
- How do you keep an agent's commits small enough to actually review?
- Only dev at my company and the AI is my only reviewer
- ChatGPT stuck on loading screen
- Your coding agent's prompt often contains old versions of files it already edited. Here's what I measured, and what I learned trying to fix it
- Chatgpt - ability to create an ai share trading system
- Anthropic just dropped the greatest advertisement for GLM ever.
- AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090
- GLM-5.3 and the Spread of Advanced Cyber Capabilities \ Anthropic
- Reflection 70B was released two years ago (September 2024)
- Deepseek Harness app is out now!!!
- [Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw
- RAM Offloading with vLLM - tcclaviger appreciation post
- Is AI Profitable Yet?
- IQuestLab/IQuest-Q1 · Hugging Face
- Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide
- nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face
- Recommended replacements for glm 4.7 flash
- Speculative reward hacking in coding agents
- Help me find a good stack for reversing an old online game client
- Qwen next 3.8 and 3.8 27b Vs Sonnet 5.5 low and Sonnet 5.5 medium.
- Do you need some extra memory on your DGX Spark?
- First few days of qwen3.8-flash-next on 4x R9700 - it's been really interesting so far
- Ornith-1.5 DFlash
- What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?
- 50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?
- I created a personality test for models, need more TESTS!!
- llama.cpp cublas error, how to uninstall/reinstall properly (Linux Mint)
- Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io
- NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
- Weekly Hiring Thread
- What exactly are we cheering for with every new model release
- I tried to poison an AI agent's memory. It worked 216 out of 216 times, and the dev shipped fixes within a month.
- Stop trying to prompt your way to agent safety. It's an access-control problem
- Need some suggestions on building AI agents for extracting the images.
- Have AI Agents growth been the best thing that happened to the world this year?
- OpenAI's Agents API won't kill agent frameworks
- At what point should an AI agent stop and ask for human approval?
- ChatGpt Dots ~ OpenClaw in new suit
- What setup step did you skip and later realize was actually important?
- What should an agent log when it decides not to act?
- Where do you enforce spend or position limits for an agent that can hit a real API: the prompt, the tool schema, or the account? Notes from watching 10 trading agents on paper money.
- Agent memory should have an expiration date
- Built a customer-support AI agent with persistent memory using Hindsight
- Self-hosted company OS, Claude Code and Codex agents in departments
- TLA+ as means to ensure heartbeats work well
- Would a separate verification agent actually solve this problem?
- What’s next after Jev? Metacache: Reasoning by Construction
- Jev Is Dead 👻
- Help needed !!
- Anyone doing content marketing for their AI automation/agent business? What's working for you?
- Github: reggaesharkk
- Ai agent
- I Built a Meeting Agent with Hindsight for Persistent Memory
- Stopped getting caught
- Dev Day i am expecting this
- Worst Dev Day.
- The rug pull was real
- LEAKED: OpenAI's revenue run rate nears $70B as enterprise sales double
- Super underwhelming DevDay...
- Where is the love for Chat, OpenAI?
- Dots isn’t avaliable in Europe
- My expectation were low but holy crap they really did get caught off guard by Anthropic.
- Dev Day schedule
- Yay or nay?
- Built a 3D fractal world explorer powered by webGL, using GPT 6 and Opus 5.5
- I'm I missing something?
- BA Computer Science, but fell in love with machine learning and AI. Just got my personal research accepted at NeurIPS as a poster. [R]
- I wrote a free, open-source book on making ML models actually fast, from silicon to agents [P]
- Functional Gradient Descent with Adaptive Representations [R]
- Advice on choosing university for PhD [D]
- NeurIPS Education Track [D]
- CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]
- Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
- Limited compute, targeting CVPR: rerun experiments for statistically strong numbers or focus on writing? [D]
- What are the trending topics in medical imaging? [D]
- Qwen3-VL 8B on a laptop vs Opus 5.5 / Sonnet 5 / GPT-5.6 on 137 messy documents: beat GPT-5.6 on tax forms, lost badly on Indian date formats[R]
- Browser demo of our Clash Royale RL environment: a 5.6k-parameter REINFORCE policy learns defensive placement against a brute-force optimum [P]
- Why not just have one less feature before softmax? [D]
- Are there any good research papers around Text clustering using LLMs [R]
- Google Deepmind Optimization Roles [D]
- OpenTrainDNN: A Browser-Based Real-Time Neural Network Visualizer.[P]
- How can I turn an industry ML project into a publication? [R]
- Quoting @joedaroo
- Quoting Muse AI Agent
- [AINews] Opus 5.5 is good at explainer videos
- Claude Code’s Next Era — Thariq Shihipar, Anthropic
- Gladys Assistant 5
- Supertake
- Timeful
- Codex Remote
- Would you pay?
- ShipHappens:
- GhostDeck
- Timeless Code
- Clink
- Tipword
- MuM
- Dina 4.5
- Xiaomi Claims an 80% Compute Cost Drop for Unreleased MiMo V3 - Geeky Gadgets
- US tech leaders are urgently calling for rules on AI – China already has them - The Conversation
- GT Economic Investigates: What does DeepSeek's doubling annualized revenue mean for China’s AI commercialization? - Global Times
- Chinese tech stocks trail US AI peers even as Huawei and DeepSeek gain ground - Startup Fortune
- US, Chinese Firms Clash Over AI Distillation Techniques - 조선일보
- China’s AI Safety Catch-Up Is Hiding in Plain Sight - The Wire China
- MiniMax M3.1 vs GPT-6 Luna vs DeepSeek V4.1 Flash: 3x Price Gap [2026] - tech-insider.org
- DeepSeek Hits $1 Billion Annual Revenue, Doubling Growth Ahead of Planned IPO - eu.36kr.com
- Expected Anthropic Sonnet 5.5 Launch Next Week Will Hopefully Reduces Costs - Geeky Gadgets
- DeepSeek Unveils New Paper: First Public Reveal of V4.1 Agent Training "Headquarters" Led by Signatory Liang Wenfeng - eu.36kr.com
- DeepSeek, Qwen, Kimi… How good are the top 10 Chinese AI players? - Telquel.ma
- Elon Musk’s AI Grok Bot can now handle banking while your Tesla FSD handles the road - Teslarati
- xAI offers to buy residents' homes in exchange for not joining class action lawsuit - FOX13 Memphis
- My Turn | My dinner with Claude and Grok - The News-Gazette
- SpaceX’s Grok Is “the Leading Client” for AI Traders on Coinbase, With 60% Share - TIKR.com
- Grok Puts the AI Assistant Between Consumers and Their Banks - PYMNTS.com
- Grok Is Everywhere: 5 Details Behind xAI's Ad Push - basenor.com
- How to Use Grok to Load a Website and Read Articles in Your Tesla - notateslaapp.com
- xAi’s models to catch up OpenAI and Anthropic by year-end - Electronics Weekly
- Grok AI Predicts XRP Could Hit $40 in 2026 With Landmark Event - Cryptonews
- Niche AI tools pose major cybersecurity risk to infrastructure operators - Cybersecurity Dive
- Tesla Owner Shows Grok AI Creating Video While FSD Drives in Rain - x.com
- The Day Your AI Doesn't Show Up for Work: How a Simultaneous Outage Exposed Business Risks - LIGA.net
- OpenCode AI Coding Agent Flaw Lets Malicious Websites Execute Code on Developer Machines - CyberSecurityNews
- OpenCode vs Codex CLI vs Gemini CLI: 103K Star Gap [2026] - tech-insider.org
- North Korean Hackers Leverage AI in Cyberattacks and Undercover Jobs - 조선일보
- Hackers Built an AI-Powered Attack Machine and Accidentally Left the Control Panel Open - CyberSecurityNews
- ゲームの敵をLLM の蒸留・強化学習・共進化で強くしてみた (いいね相当スコア: 0)
- Ollama の JSON Schema と tool calling で点検結果を受け取る — 形式と内容を分けて観察する (いいね相当スコア: 0)
- skill-creator は自分のルールをどこで破っているか: 500行の予算、全大文字、検証スクリプトと仕様のズレ (いいね相当スコア: 0)
- zvec-grepでローカルの検索性能向上を試してみる (いいね相当スコア: 0)
- AIに書かせた文書を、同じAIに採点させない。ローカルLLMだけで差し戻しが起きるか測った (いいね相当スコア: 1)
- MacBook Air M4 16GBのローカルLLMを1日で打ち切った ── 天井は12.7GB、24Bと31Bは完走しなかった (いいね相当スコア: 0)
- セミナー要点メモ: 2026/09/24 ハーネス設計入門 基礎知識の整理から実務へのステップアップ (いいね相当スコア: 0)
- Qwen Image 2.1 GGUF入門:量子化とローカル実行の始め方 (いいね相当スコア: 0)
- LLMにURLを1文字も書かせない — 内部リンクはプレースホルダで受けてDBから組み立てる (いいね相当スコア: 0)
- 「これ意味あるかな?」Claude Codeに入れたJevのプラグインを1週間で外すまで (いいね相当スコア: 2)
- Claude Code / Codexで「私のlimit、減りすぎ…?」と思ったときに見る記事 (いいね相当スコア: 3)
- 文章を書かない AI「Jev」を触って分かった、判定だけを返す API の使いどころ (いいね相当スコア: 2)
- RAG精度を自己進化させる方法。知能が安い時代に。 (いいね相当スコア: 5)
- エラーをAIに直させる自動デバッグ機能をJevを使用して作ってみた (いいね相当スコア: 2)
- 公式発表と違う「言い切り」を、生成の段階で止める (いいね相当スコア: 0)
- 自宅学習で応答30ms!軽量0.8Bの意思決定モデル「Jeff」が話題 (いいね相当スコア: 0)
- 「具体はAI、抽象は人間」──その抽象、どこから来るの? (いいね相当スコア: 0)
- すぐ陳腐化する大量データをLLMで低コスト・高速・高精度に仕分けるには? (いいね相当スコア: 2)
- local で完結する「会議の外部記憶」Exolobe の紹介 (いいね相当スコア: 0)
- 「エラーは出ないが内容が消える」——LLM抽出でハマった落とし穴と対策 (いいね相当スコア: 0)
- 自然言語処理で使われるAttentionのWeightを可視化する (いいね相当スコア: 0)
- 自然言語処理で使われるAttentionのWeightを可視化する(spaCy版) (いいね相当スコア: 0)
- 雑メモ・PyTorch 実践:DataLoader のバッチ化の仕組みとテンソル操作の注意点 (いいね相当スコア: 0)
- 欠損値の削除は行だけじゃない — 特徴量削除とペアワイズ削除の使いどころ (いいね相当スコア: 1)
- 機械学習の理論と実装をつなぐ4冊 (いいね相当スコア: 0)
- Dense Retrievalを実装する:DPR・距離尺度・ANNの設計 (いいね相当スコア: 0)
- 魚が見つからないと、お盆を測っていた ― 空振りを黙って埋めるフォールバックの話 (いいね相当スコア: 0)
- RLVRを理解する②──DatabricksでのSQL生成品質向上例 (いいね相当スコア: 1)
- 答えは合っているのに不正解? LLMベンチマークの点数はどう決まるのか (いいね相当スコア: 1)
- Jevを8,800回叩いて較正を測った ── 「確率を信じるな」は半分まちがいだった (いいね相当スコア: 0)
- 欠損値のリストワイズ削除 — dropna() の使い方と「消してはいけない」落とし穴 (いいね相当スコア: 0)
- 【学会報告】日本分析化学会第75年会に参加しました (いいね相当スコア: 6)
- レガシーPHPを壊さずリファクタリング — Claude Codeで神テーブルと激遅クエリを解体した実録 (いいね相当スコア: 0)
- VRAM超過RAM未満サイズのLLMは何とか動くのか (いいね相当スコア: 0)
- Docker ComposeでLLMアプリの開発環境と本番構成を分離する (いいね相当スコア: 0)
- 車載映像によるインフラ点検の課題その② ― 連続検出した地物をカルマンフィルタで対応付けしてGIS管理 (いいね相当スコア: 0)
- 不適切なセンサーフィードは優れたVLAモデルを台無しにする:その根拠と対策 (いいね相当スコア: 0)
- 画像生成AIのしくみ(第七話) (いいね相当スコア: 0)
- 【技術解説】Python数理ロジックと機械学習を用いた資産運用・資金管理アルゴリズムの実装ガイド (いいね相当スコア: 0)
- 口調を変えただけでは、“うちの子”にならない。相手によって顔を変えるAIキャラクターの作り方 (いいね相当スコア: 取得失敗)
- AIの利用コストを見直すために、月報生成Agentを自作し始めた (いいね相当スコア: 取得失敗)
- 近ごろの人間界…てさあ。 (いいね相当スコア: 取得失敗)
- 予防AIの規制上の地位は設計で分かれる / 欧州議会STOA予防プラットフォーム報告書 雑感 (いいね相当スコア: 取得失敗)
- VRAM12GBで始めるローカルAI駆動開発 ~ミドルレンジGPU x Bonsai2でバイブコーディングを始めよう!~ (いいね相当スコア: 取得失敗)
- 噂のローカルLLM推論エンジン「Splash」を、32GBのM6 Mac miniで動かしてみた (いいね相当スコア: 取得失敗)
- AIが増幅するサイバーリスク / 損害保険約款の射程と限度額の十分性 / Swiss Reサイバー市場報告 雑感 (いいね相当スコア: 取得失敗)
- 認知的分岐論文と委任フィードバックループ仮説 / 過度なAI依存の規律と測定の距離 雑感 (いいね相当スコア: 取得失敗)
- ローカルLLMが「メモリに入るのに遅い」のはなぜ?量子化・KVキャッシュ・生成速度を分けて理解する (いいね相当スコア: 取得失敗)
- Copilotの画像編集を人が評価していた / 人による確認の開示と写り込む第三者 雑感 (いいね相当スコア: 取得失敗)
- OpenAIの料金は2層に分かれている。公式ページで個人向け4プランとAPIはいくらなのか (いいね相当スコア: 取得失敗)
- 「OpenBot」と呼ばれる複数のオープンソースプロジェクト (いいね相当スコア: 取得失敗)
- LLM Broker/Vision BrokerからInference Brokerへの移行 (いいね相当スコア: 取得失敗)
- うちのこ達の出来上がり方(備忘録) (いいね相当スコア: 取得失敗)
- AI英会話のおすすめはどこで決まるのか。3つの層と公式の価格表から、選ぶ順番を組み立てる (いいね相当スコア: 取得失敗)
- 文章生成を捨てた判断特化型AI「JEV」について学んだことの備忘録 (いいね相当スコア: 取得失敗)
- 燃費と効率[ほぼ中×AI] (いいね相当スコア: 取得失敗)
- 【生成AIニュース+】『Claude Sonnet 5.5』『Irodori-TTS-v4-Large』『Eleven v4 / Eleven v4 Turbo』『Kling 4.0 / Kling 4.0 Flash』『Vlo 0.3』『DeepTrace』『InkDoc』『CYBER-FROST-3.8-GGUF』『UniMate』『Noct Q Anime』『Qwen-Image-2.1-Consistency-LoRA』『MiniMax-H3-FL2VA-norefiner』他 (いいね相当スコア: 取得失敗)
- 「AIを使用して執筆する」とはどういうことか~ゴンクール賞の審査除外の件で (いいね相当スコア: 取得失敗)
- AIエージェントの正体は「脳・記憶・手足」 だった!⭐Copilotは協議共同記事 (いいね相当スコア: 取得失敗)
- 広告がなくなる日:デバイスを旅したメッセージが、LLMの「食糧」になる未来 (いいね相当スコア: 取得失敗)
- 土台を変えずにコーディング性能を50%上げた——GLM-5.3が示したポストトレーニングの経済学 (いいね相当スコア: 取得失敗)
- 「生命の複雑さ」に挑んだ経験からビジネスへ。生物学の研究で培った「データに向き合う力」がキャリアを拓く
- 【STEM×ミライ応援プロジェクト|真夏の女子中高生編 #1】データの面白さを伝える!「Girls Meet STEM」オフィスツアー開催レポート
- Google Cloud、AIエージェントからの大量アクセスをPostgreSQLのプライマリDBから切り離せる「PostgreSQL for agents in AlloyDB」発表
- コーディングエージェントにGoogle Cloudの専門知識とツールを組み込む「Google Cloud Developer Plugin for AI Coding Agents」発表
- AWS、AIエージェントを自作できるツール「Strandsハーネス」をオープンソースで公開。特定のLLMに依存せず入れ替え可能、任意のコンテナ環境にデプロイ
- チームの記憶を検索可能にしたら、AIにタスクを丸ごと任せられるようになった
- Kafkaアラートの初動調査を支援するAI Alert Watcherの開発(インターンレポート)
- 2026年10月の技術系イベント予定
- Decaton Per-Key-Quotaによるプッシュ通知のスパイク制御(インターンレポート)