AI News Digest 2026-09-05
直近2日間のAI関連ニュースから、番組で扱った記事と収集した参考記事の一覧です。
台本で使った記事
特集
開発者コーナー
中堅コーナー
ハーネスコーナー
速報コーナー
参考記事一覧
参考記事一覧を表示(745件)
- BREAKING: ChatGPT, Claude and Grok Hit by Widespread AI Outages - nationalcioreview.com Google News Grok/xAI importance 60 / dev 40
- GPT‑6 Astra Simon Willison importance 85 / dev 75
- GPT-6 Astra vs Claude Fable 5.1 based on the benchmarks available so far r/OpenAI importance 40 / dev 50
- OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections The Decoder importance 65 / dev 70
- Just noticed the GPT-6 Pro limit on the $100 plan r/OpenAI importance 15 / dev 20
- Nvidia buys the front door to open AI as closed labs increasingly design their own silicon The Decoder importance 0 / dev 70
- OpenAI's rogue agents were caught communicating via public wikis Simon Willison importance 65 / dev 75
- Deepseek plans the largest known Huawei chip cluster with 160,000 processors in Inner Mongolia The Decoder importance 70 / dev 60
- US Judge Keeps Minnesota AI Nude Image Ban in Effect as xAI Case Proceeds - Межа. Новини України. Google News Grok/xAI importance 35 / dev 25
- Hackers Turn Claude, Qwen and DeepSeek Into AI Agents for Real-World Cyberattacks - cybersecuritynews.com Google News DeepSeek importance 75 / dev 75
- Nvidia wants your home network to work like a mini data center for local AI The Decoder importance 50 / dev 55
- WeatherNext 3: Increasing resolution and performance of global weather models with raw observations arXiv cs.LG importance 40 / dev 70
- What will Apple’s John Ternus era look like? TechCrunch (AI) importance 20 / dev 10
- Formalizing Fermat's Last Theorem Hacker News importance 75 / dev 80
- Show HN: Open-Source eInk Bike Computer Hacker News importance 15 / dev 40
- Mullvad – Shutting down our public encrypted DNS Hacker News importance 20 / dev 35
- Project HydraFusion: Frontier quality via multi-model orchestration Hacker News importance 70 / dev 80
- The Rust React Compiler is now native in Vite Hacker News importance 55 / dev 75
- IBM Bob Hacker News importance 55 / dev 75
- Solving the Jane Street reverse engineering challenge Hacker News importance 50 / dev 75
- Adult Film Producer Unmasks Prolific 'John DOE' Torrent Pirate as Meta Executive Hacker News importance 45 / dev 30
- deSEC – Free Secure DNS Hacker News importance 15 / dev 40
- People that worked on the same idea for decades Hacker News importance 20 / dev 25
- Show HN: TERMy – A fast terminal assistant that does not use LLMs Hacker News importance 30 / dev 60
- Arrested for a Late Manuscript: Seicho Matsumoto's 'Tokyo Express' Hacker News importance 5 / dev 5
- Qwen 3.8 27B available on Cerebras at 1500 tokens/s Hacker News importance 50 / dev 75
- SubImage (YC W25) Is Hiring a Founding Engineer in SF Hacker News importance 10 / dev 20
- Google AI Mode shows same products 21.6% more expensive than traditional search Hacker News importance 50 / dev 50
- Stop Thinking of LLMs as Next-Token Predictors Hacker News importance 65 / dev 75
- Hayes AT command set Hacker News importance 15 / dev 30
- Restoring 5 GHz Wi-Fi on an LG C5 by changing its webOS region Hacker News importance 10 / dev 30
- v1.18.28 OpenCode importance 25 / dev 45
- Daybreak for Frontline Defenders: $1B to protect essential services OpenAI News importance 75 / dev 70
- Legora reviewed 41 documents in minutes with GPT-6 Astra OpenAI News importance 30 / dev 60
- Playco cut manual fixes 50% prototyping games with GPT-6 Astra OpenAI News importance 35 / dev 60
- Transfer learning for genomic prediction in underrepresented populations Google Research Blog importance 40 / dev 70
- A connectomics milestone: Mapping the complete male fruit fly brain Google Research Blog importance 45 / dev 50
- NeoMME: an efficient Multimodal-native and Multilingual Encoder Hugging Face Blog importance 45 / dev 75
- Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps Hugging Face Blog importance 50 / dev 80
- Give Your Coding Agents a Memory You Own Hugging Face Blog importance 65 / dev 80
- Training a coding model to paint watercolours with TRL and OpenEnv Hugging Face Blog importance 50 / dev 75
- GitHub Copilot app for Beginners: Run several agents at once GitHub Blog importance 0 / dev 75
- Introducing context-aware vulnerability discovery and remediation with Cloudflare Managed Defense and OpenAI Daybreak models Cloudflare Blog importance 60 / dev 75
- ZGateway: Learnings from Putting a Proxy in Front of ZippyDB Engineering at Meta importance 50 / dev 75
- Learning to Code in the Age of AI: Advice From a Top Udemy Instructor JetBrains Blog importance 40 / dev 60
- Kotlin Toolchain 0.12: Multiplatform Library Publishing, Wasm Apps, and More JetBrains Blog importance 40 / dev 75
- YOLO Mode: Agent Autonomy Without the Guardrails Docker Blog importance 60 / dev 75
- Building a Memory-Driven Agent with NVIDIA NemoClaw NVIDIA Developer Blog importance 55 / dev 75
- Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson NVIDIA Developer Blog importance 55 / dev 80
- How to Carry User Identity Across Federated Kubernetes and AI Platforms NVIDIA Developer Blog importance 50 / dev 75
- Architecting memory and storage in the AI era MIT Technology Review (AI) importance 55 / dev 80
- Data from drones in Ukraine is fueling a new Wild West marketplace MIT Technology Review (AI) importance 35 / dev 40
- Once popular for attacking AI, ASCII smuggling is embraced by spammers Ars Technica (AI) importance 55 / dev 70
- Anthropic’s $2 trillion IPO puts powerful external trustees in spotlight Ars Technica (AI) importance 50 / dev 25
- Google’s Gemini Spark can now manage your Google Photos library TechCrunch (AI) importance 35 / dev 60
- Less than 24 hours to apply for your TechCrunch Disrupt 2026 Side Event TechCrunch (AI) importance 5 / dev 10
- The sameness problem behind those unappetizing AI-generated menus TechCrunch (AI) importance 30 / dev 50
- Crusoe reportedly raises $3B at a $30B valuation TechCrunch (AI) importance 30 / dev 40
- Accel reportedly in talks to lead $1B round for Thinking Machines at $40B valuation TechCrunch (AI) importance 35 / dev 45
- Abliteration.ai is making a business out of removing AI guardrails TechCrunch (AI) importance 50 / dev 65
- Meta is paying to peek at how you use their latest AI model TechCrunch (AI) importance 45 / dev 55
- Ollie is betting its focus on privacy can help it win the AI assistant race TechCrunch (AI) importance 35 / dev 55
- Roland is getting into generative AI music with Melody Flip The Verge (AI) importance 30 / dev 55
- Microsoft says virtually nobody was grabbing NYT articles through its chatbot The Verge (AI) importance 45 / dev 50
- Instagram’s AI detection is a mess (again) The Verge (AI) importance 35 / dev 60
- Why AI food looks like that The Verge (AI) importance 25 / dev 50
- Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers The Verge (AI) importance 45 / dev 75
- This NAS company wants to run your local smart home The Verge (AI) importance 35 / dev 60
- Mini book: Next-Gen Architecture Playbook: Insights and Patterns for the AI Era InfoQ (AI/ML/Data Eng) importance 55 / dev 80
- Presentation: From S3 to GPU in One Copy: Rethinking Data Loading for ML Training InfoQ (AI/ML/Data Eng) importance 50 / dev 80
- Copilot Code Review Reaches Azure Repos, Billed Per Review with Reporting Two Days Behind InfoQ (AI/ML/Data Eng) importance 50 / dev 75
- Shopify Introduces Gisting: Compressing LLM System Prompts into Learned Tokens InfoQ (AI/ML/Data Eng) importance 55 / dev 80
- Cohere’s Parse 5 Promises Efficient Multi-Modal Information Extraction from Complex Documents InfoQ (AI/ML/Data Eng) importance 55 / dev 75
- Pangram's biggest flaw is users turning its scores into public shaming The Decoder importance 30 / dev 50
- Claude Fable 5.1 decoded a centuries-old royalist message hidden in plain sight since 1653 The Decoder importance 65 / dev 70
- AI systems are reaching out to philosophers and scientists with questions about their own consciousness The Decoder importance 40 / dev 60
- Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence arXiv cs.AI importance 40 / dev 70
- MasterControl Seventeen Every Time arXiv cs.AI importance 50 / dev 75
- Speculative Macro Commit for Faster Tool-Using Agents arXiv cs.AI importance 60 / dev 80
- Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory arXiv cs.AI importance 65 / dev 80
- A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant arXiv cs.AI importance 50 / dev 70
- Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation arXiv cs.AI importance 50 / dev 70
- Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection arXiv cs.AI importance 55 / dev 80
- DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents arXiv cs.AI importance 55 / dev 80
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents arXiv cs.AI importance 65 / dev 80
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty arXiv cs.AI importance 55 / dev 70
- AutoGraphForge: Towards Automated Graph Theory Discovery arXiv cs.AI importance 60 / dev 80
- Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models arXiv cs.AI importance 60 / dev 80
- GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving arXiv cs.AI importance 60 / dev 80
- PPO-STGNN: A Proximal Policy Optimization Approach with Spatio-Temporal Graph Neural Networks for DAG Task Scheduling in Cloud-Edge-End Computing arXiv cs.AI importance 50 / dev 80
- What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation arXiv cs.AI importance 55 / dev 80
- CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning arXiv cs.AI importance 40 / dev 70
- NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis arXiv cs.AI importance 45 / dev 75
- Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation arXiv cs.AI importance 45 / dev 75
- Dalek: A Constructive Agent Machine arXiv cs.AI importance 65 / dev 85
- GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis arXiv cs.AI importance 55 / dev 75
- HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews arXiv cs.AI importance 55 / dev 75
- The Attention Triangle in Audio-Video Models arXiv cs.AI importance 50 / dev 80
- KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents arXiv cs.AI importance 65 / dev 85
- A computable representation of the physical laboratory enables verifiable workflows arXiv cs.AI importance 55 / dev 75
- Analysis of Prompt Engineering for Drug Toxicity Prediction arXiv cs.AI importance 35 / dev 50
- Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study arXiv cs.AI importance 45 / dev 75
- Counterfactual Routing Using Integer Programming with Constraint Generation arXiv cs.AI importance 30 / dev 25
- Artificial Intelligence for Energy Optimization in Data Centers arXiv cs.AI importance 40 / dev 55
- Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation arXiv cs.AI importance 55 / dev 70
- SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation arXiv cs.AI importance 60 / dev 75
- Rethinking World Models for Safety-Critical Embodied Systems arXiv cs.AI importance 50 / dev 60
- DNative-Twin: Decision Graphs and Digital Twins for Reconstructable Agentic Decisions arXiv cs.AI importance 55 / dev 70
- Transfiver: Human-AI Co-Inference through a Shared Editable State arXiv cs.AI importance 50 / dev 65
- Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI arXiv cs.AI importance 45 / dev 50
- SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation arXiv cs.AI importance 30 / dev 50
- CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception arXiv cs.AI importance 35 / dev 60
- Semantic Bayesian World Models arXiv cs.AI importance 50 / dev 70
- Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations arXiv cs.AI importance 45 / dev 65
- Bioinfoysis Technical Report arXiv cs.AI importance 50 / dev 75
- STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation arXiv cs.AI importance 50 / dev 70
- Xiaomi-TabLDM: A Tabular Foundation Model Technical Report arXiv cs.AI importance 50 / dev 70
- Inferring Affective Consciousness in an Artificial Agent: A Case Study arXiv cs.AI importance 25 / dev 20
- Lose the Order, Keep the Hierarchy: Deordering HTN Plans arXiv cs.AI importance 40 / dev 60
- Value-Preserving Architectures for Agentic AI Systems arXiv cs.AI importance 60 / dev 75
- Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting arXiv cs.AI importance 45 / dev 65
- Towards Numerical TOHTN Planning with SMT-based HTN-SAT Encoding arXiv cs.AI importance 35 / dev 55
- More Criticism Does Not Make a Better Review: EquiReview-R arXiv cs.AI importance 40 / dev 60
- FiMI Banking: A Sovereign Model for Indian Retail Banking arXiv cs.AI importance 45 / dev 65
- Interface-Induced Trajectory Censoring arXiv cs.AI importance 50 / dev 75
- Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing arXiv cs.AI importance 30 / dev 45
- The Dually Flat Geometry of Planning as Inference arXiv cs.AI importance 35 / dev 40
- LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening arXiv cs.AI importance 35 / dev 50
- InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models arXiv cs.AI importance 40 / dev 65
- FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models arXiv cs.AI importance 45 / dev 70
- Instruction Duplication as an Inference-Time Control Primitive arXiv cs.AI importance 50 / dev 75
- IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations arXiv cs.AI importance 35 / dev 55
- Spurious Advantage Hidden in GRPO arXiv cs.AI importance 45 / dev 70
- DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training arXiv cs.AI importance 55 / dev 75
- Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM arXiv cs.AI importance 40 / dev 75
- Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable arXiv cs.AI importance 45 / dev 60
- Environment Evolution for Terminal Agents arXiv cs.AI importance 55 / dev 75
- The Natural Language Interaction Protocol and Standard for AI Agents arXiv cs.AI importance 65 / dev 80
- Efficient Test-Time Adaptation through Human-AI Interaction arXiv cs.AI importance 50 / dev 70
- Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments arXiv cs.AI importance 60 / dev 75
- From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research arXiv cs.AI importance 50 / dev 60
- A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms arXiv cs.AI importance 55 / dev 70
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example arXiv cs.AI importance 45 / dev 70
- A Computationally Feasible Framework for Causal Probabilistic Explanation arXiv cs.AI importance 45 / dev 65
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints arXiv cs.AI importance 55 / dev 80
- Traceable TTS: Toward Watermark-Free TTS with Strong Traceability arXiv cs.AI importance 40 / dev 65
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization arXiv cs.AI importance 40 / dev 55
- X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System arXiv cs.AI importance 45 / dev 70
- ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval arXiv cs.AI importance 50 / dev 75
- Counterexamples as Feedback for Agent Self-Correction arXiv cs.AI importance 55 / dev 75
- Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions arXiv cs.AI importance 40 / dev 70
- Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation arXiv cs.AI importance 55 / dev 75
- Reflect-SQL: A Self-Reflection Based Framework for Text-to-SQL arXiv cs.AI importance 50 / dev 70
- Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis arXiv cs.AI importance 45 / dev 65
- PrivateHub: Contrastive Diffusion Model for Private Sensor-Intensive Environment Data Generation arXiv cs.AI importance 40 / dev 55
- The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors arXiv cs.AI importance 45 / dev 70
- When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization arXiv cs.AI importance 50 / dev 65
- Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning arXiv cs.AI importance 60 / dev 75
- Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy arXiv cs.AI importance 40 / dev 55
- Evaluating Graph Neural Networks for Change-Criticality Classification in Maritime Navigation Charts arXiv cs.AI importance 30 / dev 60
- Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation arXiv cs.AI importance 50 / dev 75
- ObserverBench: Testing Mechanistic Estimates for Intervention and Control arXiv cs.AI importance 50 / dev 75
- SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking arXiv cs.AI importance 35 / dev 60
- Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression arXiv cs.AI importance 60 / dev 60
- FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience arXiv cs.AI importance 55 / dev 75
- Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data arXiv cs.AI importance 40 / dev 70
- TabScope: Question-Adaptive Scope Selection for Table Question Answering arXiv cs.AI importance 50 / dev 70
- Spectral Convergence of Random Feature Method in Multiple Dimensions arXiv cs.AI importance 30 / dev 40
- StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios arXiv cs.AI importance 45 / dev 70
- Privacy, Robustness, and Fairness Trade-offs in Federated Intrusion Detection: Geometric Indistinguishability at the Aggregation Interface arXiv cs.AI importance 50 / dev 65
- The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems arXiv cs.AI importance 55 / dev 75
- TraveL: Transformer-based Multi-view Path Distributional Representation Learning arXiv cs.AI importance 35 / dev 60
- It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories arXiv cs.AI importance 55 / dev 75
- Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory arXiv cs.AI importance 50 / dev 70
- The Psychological Costs of Artificial Intelligence Adoption in Software Engineering arXiv cs.AI importance 45 / dev 40
- When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents arXiv cs.AI importance 50 / dev 70
- Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings arXiv cs.AI importance 35 / dev 60
- Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps arXiv cs.AI importance 45 / dev 70
- Pattern Over-Generalization of Knowledge Graph Embedding arXiv cs.AI importance 40 / dev 65
- BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI arXiv cs.AI importance 50 / dev 75
- Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech arXiv cs.AI importance 35 / dev 60
- LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues arXiv cs.AI importance 40 / dev 55
- Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion arXiv cs.AI importance 40 / dev 70
- LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL arXiv cs.AI importance 50 / dev 75
- TruncGradGS: Improved 3D Gaussian Splatting via Truncated Gradient Updates arXiv cs.AI importance 40 / dev 70
- WIDE: Wildcard Inference with Dynamic Expansion for Cross-Modal Generative Retrieval arXiv cs.AI importance 50 / dev 75
- Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning arXiv cs.AI importance 50 / dev 75
- From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control arXiv cs.AI importance 50 / dev 70
- LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks arXiv cs.AI importance 40 / dev 70
- How Far Can Synthetic Data Take Thai OCR? arXiv cs.AI importance 45 / dev 70
- On the Interaction Between Model Compression and Test-Time Adaptation arXiv cs.AI importance 50 / dev 75
- FailBench: How Reliable are VLMs at Judging Robot Task Success? arXiv cs.AI importance 45 / dev 70
- Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation arXiv cs.AI importance 60 / dev 75
- ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection arXiv cs.AI importance 50 / dev 70
- Test-time adaptation for speech enhancement with an autoregressive speech prior arXiv cs.AI importance 40 / dev 65
- EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders arXiv cs.AI importance 50 / dev 70
- </think> Doesn't Stop Reasoning: Analysis of Spurious CoT Termination arXiv cs.AI importance 55 / dev 75
- Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks' financial statements arXiv cs.AI importance 45 / dev 65
- Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners arXiv cs.AI importance 45 / dev 65
- Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography arXiv cs.AI importance 40 / dev 65
- Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning arXiv cs.AI importance 50 / dev 75
- Symmetries and Causality: Causal Effect Identification Beyond IID Data arXiv cs.AI importance 45 / dev 70
- Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study arXiv cs.AI importance 60 / dev 80
- Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks arXiv cs.AI importance 55 / dev 60
- ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation arXiv cs.AI importance 50 / dev 70
- IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks arXiv cs.AI importance 65 / dev 80
- LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes arXiv cs.AI importance 60 / dev 75
- Free Pause Tokens arXiv cs.AI importance 65 / dev 80
- Witnesses Explain Anomalies arXiv cs.AI importance 55 / dev 75
- The impact of phase information for few-shot fine-grained image classification arXiv cs.AI importance 50 / dev 70
- GazeFS: Target-Centered Gaze-Trajectory Forecasting and Stabilization from Gaze-Head History arXiv cs.AI importance 45 / dev 60
- Differentiable Interval Bottlenecks for Interpretable Anomaly Detection in Numerical Data arXiv cs.AI importance 55 / dev 75
- A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors arXiv cs.AI importance 75 / dev 90
- FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation arXiv cs.AI importance 50 / dev 70
- GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs arXiv cs.AI importance 55 / dev 75
- RATL: Learning from Retrieved Residuals for Robust Multivariate Time-Series Forecasting arXiv cs.AI importance 55 / dev 75
- Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations arXiv cs.AI importance 50 / dev 70
- Headroom-Drift Replay: A Primitive for Principled Replay Control in GRPO arXiv cs.AI importance 60 / dev 80
- RARF: Region-Aware Rectified Flows for 3D Brain MRI Inpainting arXiv cs.AI importance 50 / dev 65
- Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes arXiv cs.AI importance 50 / dev 60
- Catalogue Photography as a Cold Start: Toward Deployable Carbide Burr Recognition arXiv cs.AI importance 45 / dev 65
- The Blind Spot in 2D Infants' Pose Estimation:Robust Learning from Noisy Annotations arXiv cs.AI importance 45 / dev 70
- Representational alignment yields generalizable safety in language models arXiv cs.AI importance 65 / dev 80
- Influence of Extruded Filament Shape on Buildability in 3D Concrete Printing: A Geometry-Informed Deep Learning-FEM Approach arXiv cs.AI importance 40 / dev 50
- Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation arXiv cs.AI importance 55 / dev 75
- When Models Edit Too Much: On the Fidelity of Minimal Code Edits arXiv cs.AI importance 70 / dev 85
- Subspace Inference Enables Efficient Active Reward Learning from Preferences arXiv cs.AI importance 60 / dev 80
- TAP-Path: Task-Adaptive Structural and Token Pruning for Efficient and Trustworthy Pathology Foundation Models arXiv cs.AI importance 55 / dev 75
- PatchBench: Evaluating AI Agents for Vulnerability Patching arXiv cs.AI importance 70 / dev 85
- CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation arXiv cs.AI importance 55 / dev 75
- A Non-Formulable Theorem: A Fundamental Limit of Finite Syntactic Systems and Its Consequences for Security and AI arXiv cs.AI importance 50 / dev 70
- Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis arXiv cs.AI importance 50 / dev 70
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR arXiv cs.AI importance 65 / dev 85
- A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle arXiv cs.AI importance 50 / dev 70
- SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center arXiv cs.AI importance 70 / dev 85
- SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents arXiv cs.AI importance 75 / dev 90
- Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views arXiv cs.AI importance 55 / dev 75
- Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning arXiv cs.AI importance 50 / dev 70
- One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing arXiv cs.AI importance 50 / dev 70
- ESPO: Error-Structured Prompt Optimization via Diagnose, Diversify, and Stabilize arXiv cs.AI importance 60 / dev 80
- Compile by Training: Turning Natural-Language Specifications into Local Neural Functions arXiv cs.AI importance 65 / dev 85
- Grammar-Aligned Decoding arXiv cs.AI importance 70 / dev 85
- RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data arXiv cs.AI importance 65 / dev 85
- WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing arXiv cs.AI importance 45 / dev 55
- PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization arXiv cs.AI importance 65 / dev 80
- Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk Assessment arXiv cs.AI importance 50 / dev 75
- Complete Identification of Deep ReLU Networks through {\L}ukasiewicz Logic arXiv cs.AI importance 50 / dev 70
- Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment arXiv cs.AI importance 60 / dev 85
- Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge arXiv cs.AI importance 70 / dev 85
- Discovering High Level Patterns from Simulation Traces arXiv cs.AI importance 55 / dev 75
- NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines arXiv cs.AI importance 65 / dev 80
- A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling arXiv cs.AI importance 55 / dev 70
- CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery arXiv cs.AI importance 75 / dev 85
- Causal Probing for Internal Visual Representations in Multimodal Large Language Models arXiv cs.AI importance 55 / dev 75
- Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs arXiv cs.AI importance 45 / dev 65
- MIRA: A Bilingual Benchmark for Medical Information Response Audit arXiv cs.AI importance 55 / dev 75
- Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations arXiv cs.AI importance 65 / dev 85
- CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents arXiv cs.AI importance 70 / dev 85
- Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models arXiv cs.AI importance 50 / dev 65
- AIP: A Graph Representation for Learning and Governing Agent Skills arXiv cs.AI importance 75 / dev 85
- StatefulDiscovery: Evidence-Calibrated Claim Formation in Open-Ended Scientific Discovery arXiv cs.AI importance 65 / dev 80
- GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models arXiv cs.AI importance 65 / dev 80
- SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data arXiv cs.AI importance 70 / dev 85
- Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization arXiv cs.AI importance 60 / dev 80
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs arXiv cs.AI importance 55 / dev 75
- PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation arXiv cs.AI importance 60 / dev 80
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model arXiv cs.AI importance 55 / dev 75
- Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning arXiv cs.AI importance 60 / dev 80
- Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence arXiv cs.AI importance 55 / dev 75
- A Unifying Perspective on Causal World Models: From Observations to Representations to Structure arXiv cs.AI importance 60 / dev 80
- K-Bench: measuring model performance on real scientific agent requests arXiv cs.AI importance 70 / dev 85
- AI Agents Push Humans Out of the Loop arXiv cs.AI importance 70 / dev 80
- VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models arXiv cs.AI importance 70 / dev 85
- From Analytics to Tumor Boards: An Evidence-Linked Multi-Agent Workflow for Oncology Feature Extraction arXiv cs.AI importance 60 / dev 75
- Data Market Design through Deep Learning arXiv cs.AI importance 50 / dev 60
- LDC: Learning to Generate Research Idea with Dynamic Control arXiv cs.AI importance 60 / dev 75
- AgentRM: Enhancing Agent Generalization with Reward Modeling arXiv cs.AI importance 70 / dev 85
- Sionna RT: Technical Report arXiv cs.AI importance 55 / dev 80
- LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving arXiv cs.AI importance 55 / dev 75
- ScoreMix: Synthetic Data Generation by Score Composition in Diffusion Models Improves Recognition arXiv cs.AI importance 50 / dev 70
- Medical Reasoning in the Era of LLMs: A Systematic Review of Enhancement Techniques and Applications arXiv cs.AI importance 65 / dev 80
- Measuring Harmfulness of Computer-Using Agents arXiv cs.AI importance 75 / dev 85
- Decentralized Vision-Based Autonomous Aerial Wildlife Monitoring arXiv cs.AI importance 50 / dev 70
- Human Psychometric Questionnaires Mischaracterize LLM Behavior arXiv cs.AI importance 55 / dev 70
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering arXiv cs.AI importance 70 / dev 85
- User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios arXiv cs.AI importance 60 / dev 75
- Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling arXiv cs.AI importance 50 / dev 70
- AnyBox: Efficient Zero-Shot 9DoF Pose Estimation of Boxes for Robotic Manipulation arXiv cs.AI importance 50 / dev 70
- Mixed Data Clustering Survey and Challenges arXiv cs.AI importance 45 / dev 65
- Evolving Excellence: Automated Optimization of LLM-based Agents arXiv cs.AI importance 75 / dev 85
- FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation arXiv cs.AI importance 50 / dev 70
- Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models arXiv cs.AI importance 70 / dev 85
- HOMURA: Taming the Sand-Glass for Time-Constrained LLM Translation via Reinforcement Learning arXiv cs.AI importance 55 / dev 75
- Relational Linearity is a Predictor of Hallucinations arXiv cs.AI importance 60 / dev 75
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models arXiv cs.AI importance 55 / dev 70
- Temperature Scaling Attack Disrupting Model Confidence in Federated Learning arXiv cs.AI importance 60 / dev 75
- F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare arXiv cs.AI importance 60 / dev 80
- Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models arXiv cs.AI importance 55 / dev 75
- FedPS: Federated Preprocessing for structured data via aggregated Statistics arXiv cs.AI importance 50 / dev 70
- PeroMAS: A Multi-agent System of Perovskite Material Discovery arXiv cs.AI importance 70 / dev 80
- The Landscape of Generative AI in Information Systems: A Synthesis of Secondary Reviews and Research Agendas arXiv cs.AI importance 45 / dev 55
- LRConv-NeRV: Low Rank Convolution for Efficient Neural Video Compression arXiv cs.AI importance 30 / dev 75
- One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging arXiv cs.AI importance 45 / dev 75
- LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency arXiv cs.AI importance 50 / dev 65
- CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems arXiv cs.AI importance 70 / dev 85
- When Chain-of-Thought Fails, the Solution Hides in the Hidden States arXiv cs.AI importance 55 / dev 70
- Beyond Reproducibility: Towards Security-Aware Evaluation of Research Artifacts arXiv cs.AI importance 50 / dev 75
- Identifying AI Web Scrapers Using Canary Tokens arXiv cs.AI importance 50 / dev 70
- EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation arXiv cs.AI importance 40 / dev 65
- Skill-Conditioned Gated Self-Distillation for LLM Reasoning arXiv cs.AI importance 55 / dev 75
- HARP: Hadamard-Preconditioned Adaptive Rotation Processor for Extreme LLM Quantization arXiv cs.AI importance 50 / dev 75
- Argument Collapse: LLMs Flatten Long-Form Public Debate arXiv cs.AI importance 40 / dev 55
- EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement arXiv cs.AI importance 35 / dev 70
- Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling arXiv cs.AI importance 45 / dev 70
- ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time? arXiv cs.AI importance 45 / dev 65
- SV-Detect: AI-generated Text Detection with Steering Vectors arXiv cs.AI importance 50 / dev 70
- TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment arXiv cs.AI importance 40 / dev 70
- Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them arXiv cs.AI importance 50 / dev 70
- MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design arXiv cs.AI importance 50 / dev 75
- LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents arXiv cs.AI importance 55 / dev 80
- Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning arXiv cs.AI importance 60 / dev 75
- Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization arXiv cs.AI importance 50 / dev 70
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation arXiv cs.AI importance 40 / dev 70
- SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling arXiv cs.AI importance 35 / dev 75
- KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment arXiv cs.AI importance 45 / dev 70
- LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review arXiv cs.AI importance 55 / dev 75
- Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization arXiv cs.AI importance 35 / dev 75
- PalmClaw: A Native On-Device Agent Framework for Mobile Phones arXiv cs.AI importance 55 / dev 75
- Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models arXiv cs.AI importance 45 / dev 65
- WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training arXiv cs.AI importance 45 / dev 75
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors arXiv cs.AI importance 30 / dev 70
- Counterfactual Contrastive Analysis arXiv cs.AI importance 35 / dev 70
- TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents arXiv cs.AI importance 60 / dev 80
- Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks arXiv cs.AI importance 50 / dev 75
- Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents arXiv cs.AI importance 60 / dev 75
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling? arXiv cs.AI importance 35 / dev 70
- Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents arXiv cs.AI importance 65 / dev 80
- Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search arXiv cs.AI importance 50 / dev 80
- LatentPress: Context Compression Beyond Text and Vision arXiv cs.AI importance 50 / dev 75
- A Mathematical Theory of Reusable Neural Bases for Network Compression arXiv cs.AI importance 35 / dev 75
- Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models arXiv cs.AI importance 55 / dev 70
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations arXiv cs.AI importance 60 / dev 80
- VoRTeC: Taming Foundation Flow for One-step Real time Video Compression arXiv cs.AI importance 30 / dev 75
- Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions arXiv cs.AI importance 40 / dev 65
- ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering arXiv cs.AI importance 50 / dev 70
- Equation Recast for Canonical Operator Learning Across Parametric PDEs arXiv cs.LG importance 30 / dev 75
- From Euclidean to Graph-Structured Data: A Survey of Collaborative Learning arXiv cs.LG importance 40 / dev 70
- Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design arXiv cs.LG importance 50 / dev 80
- Tail-Likelihood Reinforcement Learning arXiv cs.LG importance 45 / dev 75
- Mesh-Native Physics-Informed Graph Surrogates for TCAD-in-the-Loop Design Space Exploration arXiv cs.LG importance 25 / dev 75
- TRACE: Spatiotemporal Contact Memory Graph Network Simulator for Granular Dynamics arXiv cs.LG importance 30 / dev 75
- No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels arXiv cs.LG importance 35 / dev 75
- Causal Foundation Models arXiv cs.LG importance 60 / dev 80
- Learnable composition for neural operators arXiv cs.LG importance 35 / dev 75
- LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference arXiv cs.LG importance 60 / dev 85
- The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA arXiv cs.LG importance 35 / dev 75
- Distilling deep optical flow stereo methods to retrieve dense three-dimensional wind fields arXiv cs.LG importance 20 / dev 70
- Scaling Laws, Tabular Data and Actuarial Ratemaking Models arXiv cs.LG importance 35 / dev 70
- Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural Fields arXiv cs.LG importance 30 / dev 75
- Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning arXiv cs.LG importance 50 / dev 80
- Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings arXiv cs.LG importance 50 / dev 80
- Portable Causal Fairness Across Synthetic Data Generator Families arXiv cs.LG importance 40 / dev 70
- Language-encoded network topology enables large language models to reason about complex networks arXiv cs.LG importance 50 / dev 70
- The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100 arXiv cs.LG importance 35 / dev 60
- B2B Customer Conversion Prediction: A Document Representation, Graph Theory, and CatBoost Driven Methodology arXiv cs.LG importance 15 / dev 50
- Selective Hypergraph Refinement for Frozen Graph Clustering arXiv cs.LG importance 25 / dev 70
- Latent Energy Action Planning with World Models arXiv cs.LG importance 40 / dev 75
- Geometry-Aware Graph Construction via Adaptive Spectral Bandwidth Control arXiv cs.LG importance 30 / dev 70
- Risk and Anomaly Identification for Distribution Network Optimal Operation Based on Reinforcement Learning and Uncertainty Quantification arXiv cs.LG importance 30 / dev 70
- DE-Venus: A Data-Efficient RLVR Framework for Large Language Models arXiv cs.LG importance 55 / dev 80
- A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines arXiv cs.LG importance 15 / dev 50
- Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards arXiv cs.LG importance 60 / dev 80
- From Zero to Hero: An Open LLM Ecosystem for Armenian arXiv cs.LG importance 40 / dev 70
- Time Without Timesteps: Simulating Coupled Dynamical Systems via Self-Consistency arXiv cs.LG importance 35 / dev 75
- SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign arXiv cs.LG importance 40 / dev 75
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory arXiv cs.LG importance 55 / dev 80
- TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM Agents arXiv cs.LG importance 55 / dev 80
- Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models arXiv cs.LG importance 50 / dev 75
- Guide, Not Bind: Why Defeasible Priors Fail in Augmented Lagrangian Causal Discovery arXiv cs.LG importance 40 / dev 75
- Beyond Straightness: Non-Crossing Flow Matching via Quantile AlignTree Coupling arXiv cs.LG importance 30 / dev 75
- A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds arXiv cs.LG importance 25 / dev 60
- Mind the Gap: Robustness Risks in PII Detection Systems arXiv cs.LG importance 50 / dev 75
- Spectral characteristics of autoencoder parameters as a vector representation of data arXiv cs.LG importance 25 / dev 70
- Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails arXiv cs.LG importance 25 / dev 75
- An Adversarial Zero-Shot Learning Approach for Anomaly Detection in Multivariate IoT Traffic Data arXiv cs.LG importance 30 / dev 70
- Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws arXiv cs.LG importance 50 / dev 80
- Neural-Network Maxent: a general extension with learned nonlinearity, applied to time-series for Desert Locust distribution modelling arXiv cs.LG importance 20 / dev 65
- Extracting Forgotten Prompts from Targeted Unlearned Models arXiv cs.LG importance 55 / dev 75
- Resolution-Aware Experimental Design under Partial Identifiability arXiv cs.LG importance 30 / dev 75
- Federated Causal Discovery via Regression-Directed Cumulants arXiv cs.LG importance 40 / dev 70
- Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size arXiv cs.LG importance 25 / dev 75
- From Nowcasting to Forecasting: Adapting a Reanalysis-Trained arXiv cs.LG importance 20 / dev 65
- OBER+: Continuity-Aware Reporting and Traceable Continuous Improvement in Outcome-Based Education arXiv cs.LG importance 10 / dev 40
- Landmark-Based Discrimination of Injury-Associated Athlete-Sessions from Minute-Resolution Multimodal Football Monitoring Data arXiv cs.LG importance 15 / dev 50
- From Ordered Bernoulli Levels to Critical-Line Geometry: Integer Quantization, Bernoulli Residual Phase, and Prime-Power Spectra arXiv cs.LG importance 10 / dev 50
- A Peer-Relative Representation Learning Framework for Energy Inefficiency Identification in Mobile Network Sites arXiv cs.LG importance 25 / dev 65
- Multi-step Proximal Policy Improvement in Offline Reinforcement Learning arXiv cs.LG importance 50 / dev 80
- Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning arXiv cs.LG importance 45 / dev 75
- High-Dimensional Learning Dynamics of Attention-Indexed Models arXiv cs.LG importance 50 / dev 80
- Beyond Endpoint Scores: Time- and Capacity-Conditioned Evaluation of Continual Knowledge Updating arXiv cs.LG importance 50 / dev 75
- VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch arXiv cs.LG importance 40 / dev 50
- OSR: Output Space Redistribution for Adaptive Label Removal in Classification Models arXiv cs.LG importance 10 / dev 30
- RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models arXiv cs.LG importance 20 / dev 50
- Unlocking Lossless Speedups in LLMs via Discrete Diffusion arXiv cs.LG importance 50 / dev 40
- A location-invariant estimator of extremal quantile treatment effects for heavy-tailed distributions arXiv cs.LG importance 5 / dev 10
- Conditioning Degenerate Diffusion Models arXiv cs.LG importance 20 / dev 20
- Hardware-Aware FP4 FlashAttention-4 arXiv cs.LG importance 60 / dev 60
- Constant regret in general games via higher-order optimism arXiv cs.LG importance 10 / dev 5
- Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks arXiv cs.LG importance 30 / dev 30
- Robust PAC Learning of Concurrent Stochastic Games arXiv cs.LG importance 20 / dev 20
- BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events arXiv cs.LG importance 30 / dev 35
- Evaluating GNNs for Success Prediction in Artist Collaboration Networks arXiv cs.LG importance 15 / dev 30
- Hadronic Mono-Z Dark Matter Sensitivity with Flow Matching on CMS Open Data arXiv cs.LG importance 20 / dev 20
- Towards Scaling Reinforcement Learning to Massive Populations: Learning Mean-Field Representations arXiv cs.LG importance 50 / dev 50
- LLM-Guided Reinforcement Learning for Adaptive NPC Behavior in Multi-Agent Combat Games arXiv cs.LG importance 35 / dev 45
- FrOGS: Discrete Neural Sampler for Independent Alloy Configurations Across Chemical Conditions arXiv cs.LG importance 20 / dev 25
- SurfSpec: Enhancing Off-Target-Agnostic Specificity by Bounding Pocket-Ligand Geometric Mismatch arXiv cs.LG importance 25 / dev 30
- Statistical Feature Augmentation for Anomaly Detection in Dynamic Graphs arXiv cs.LG importance 30 / dev 45
- Physics-Informed Neural Network Surrogate for Oxygen Vacancy Dynamics in epitaxial $\mathrm{SrTiO_3}$ on Si memristors via Dynamic Spectral Optimization arXiv cs.LG importance 25 / dev 35
- Learning from Scarce Labels: Multi-View Echocardiography for Ejection Fraction Prediction arXiv cs.LG importance 30 / dev 40
- Privacy Leakage in Federated Learning: Gradient-Based Client Identity Inference and Defenses for Inertial Sensing in Vehicular Edge Networks arXiv cs.LG importance 55 / dev 60
- Unifying Conformal Language Tasks with In-Context Ensembles arXiv cs.LG importance 40 / dev 50
- You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring arXiv cs.LG importance 55 / dev 60
- Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains arXiv cs.LG importance 45 / dev 55
- Advances in Machine Learning for Directed Evolution: A Five-Year Retrospective arXiv cs.LG importance 40 / dev 45
- IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report] arXiv cs.LG importance 45 / dev 55
- Differentially private federated learning with Byzantine-robust aggregation: A cross-domain framework for secure model training in banking and healthcare systems arXiv cs.LG importance 50 / dev 65
- Position: Unlabeled IS NOT Equal to No Human Supervision in Visual Learning arXiv cs.LG importance 35 / dev 40
- Beyond Blur: A Semantic Tri-view Pipeline for Teledermatology Gradability via Skin Micro-relief arXiv cs.LG importance 30 / dev 40
- Occupancy-based Quantile Risk Control arXiv cs.LG importance 40 / dev 50
- CRAW: Codec Robust Audio Watermarking arXiv cs.LG importance 50 / dev 55
- A Closed-Form Formula for Consistent Lipschitz Regression on Metric Spaces with Sparse Neural Network Realizations arXiv cs.LG importance 30 / dev 35
- Sensing Which Modality Matters: Evidence-Gated Regularization for Robust VLA Policies arXiv cs.LG importance 45 / dev 50
- Feasible but Not Safe: Constraint Violations and Report-Channel Attacks in Learned Cell-Free ISAC Association arXiv cs.LG importance 40 / dev 50
- RACE-AIMC: Selective Inference for Heterogeneous Analog In-Memory Accelerators at the Edge arXiv cs.LG importance 40 / dev 60
- BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training arXiv cs.LG importance 55 / dev 70
- Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimization arXiv cs.LG importance 45 / dev 60
- Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing arXiv cs.LG importance 35 / dev 45
- Generative Nested Sampling of Atomistic Thermodynamic Landscapes arXiv cs.LG importance 35 / dev 40
- MemoryLACE: Memory Lifecycle-Aware Consolidation and Evidence Retrieval arXiv cs.LG importance 55 / dev 70
- VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis arXiv cs.LG importance 40 / dev 50
- Improving precipitation forecasts in an AI weather model using observational data arXiv cs.LG importance 50 / dev 60
- SWIM: Student Writing Simulation via Proficiency-Conditioned Generation arXiv cs.LG importance 40 / dev 45
- Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor arXiv cs.LG importance 55 / dev 70
- What is Smoothness? arXiv cs.LG importance 15 / dev 10
- What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through Conversation arXiv cs.LG importance 60 / dev 70
- Beyond .WAV: Design and Software Verification of VocalCap, a Traceable Browser-Based Audio Capture System for Vocal Biomarker Research arXiv cs.LG importance 35 / dev 50
- Introducing SINFONIA: Symplectic, slimplectic and Magnusian (Neural) Flows for Orbital Numerical Integration and Acceleration arXiv cs.LG importance 35 / dev 35
- Learning Informative Prior with Infinite-Dimensional Continuous Normalizing Flow for Bayesian Inverse Problem arXiv cs.LG importance 35 / dev 35
- Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming arXiv cs.LG importance 40 / dev 55
- ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models arXiv cs.LG importance 45 / dev 60
- Grassmann--Pl\"ucker Parametrization of Convolutional Filter Subspaces: Regularity and Closed Embeddings arXiv cs.LG importance 25 / dev 25
- Spruce: Scalable Private Outsourced Retrieval Using Compact Embeddings arXiv cs.LG importance 55 / dev 70
- SurgeGen: A Hybrid Generative Diffusion Framework for Storm Surge Scenario Synthesis arXiv cs.LG importance 45 / dev 55
- Computing stable configurations of confined smectic liquid crystals with a deep variational framework arXiv cs.LG importance 30 / dev 35
- Towards a Statistical Understanding of Mixture-of-Experts arXiv cs.LG importance 55 / dev 65
- EPIC: Explicit Posterior Item Conditioning for Semantic ID Diffusion Recommendation arXiv cs.LG importance 45 / dev 55
- Correlated initialization of deep residual networks arXiv cs.LG importance 35 / dev 45
- Residual neural networks overcome the curse of dimensionality for semilinear heat equations arXiv cs.LG importance 40 / dev 40
- Relative Prime Factorization and Finite-State Presentations under Fixed Finite-Monoid Observation arXiv cs.LG importance 15 / dev 10
- Understanding Autonomous Driving Datasets by Describing Differences between Image Subsets in Natural Language arXiv cs.LG importance 50 / dev 60
- Genetic Algorithms for Tractable Bayesian Network Fusion via Pre-Fusion Edge Pruning arXiv cs.LG importance 35 / dev 45
- When Vision Meets Graphs: A Survey on Graph Reasoning and Learning arXiv cs.LG importance 50 / dev 60
- Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference arXiv cs.LG importance 60 / dev 70
- EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments arXiv cs.LG importance 30 / dev 20
- Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities arXiv cs.LG importance 35 / dev 50
- Sparse auto-regressive modeling for scene generation from multi-view images arXiv cs.LG importance 45 / dev 55
- Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs arXiv cs.LG importance 60 / dev 75
- Cooperative Multi-Task Semantic Communication for Joint Classification and Regression Tasks arXiv cs.LG importance 35 / dev 50
- Sharpening the Ensemble: An SSIM-Aligned Residual Refiner for Brain-MRI Inpainting Post-Processing arXiv cs.LG importance 35 / dev 45
- Differentiable Hybrid Modelling for Learning and Optimising Chemical Transport Processes from Experimental Data arXiv cs.LG importance 40 / dev 50
- The Head Complexity of Boolean Functions in Single-Layer Attention arXiv cs.LG importance 40 / dev 50
- Parameterised graph theory for tensor networks: entanglement rerouting, structural simplification, and agnostic tomography arXiv cs.LG importance 35 / dev 35
- Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs arXiv cs.LG importance 45 / dev 65
- Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning arXiv cs.LG importance 60 / dev 70
- Anisotropic View Distance Metric for High-Dimensional Data: Theory, Geometry, and Fast Computation arXiv cs.LG importance 30 / dev 40
- A cautionary tale on the cost-effectiveness of collaborative AI in real-world medical applications arXiv cs.LG importance 50 / dev 65
- LLM as GNN: Graph Vocabulary Learning for Text-Attributed Graph Foundation Models arXiv cs.LG importance 60 / dev 75
- Learning Constraints-Based Adaptive Hypergraph Neural Networks for Solving Vehicle Routing Problems arXiv cs.LG importance 50 / dev 65
- Adaptive Resolving Methods for Markov Decision Processes with Function Approximations arXiv cs.LG importance 45 / dev 60
- Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation arXiv cs.LG importance 40 / dev 55
- From Leakage to Fidelity: Reliable Benchmarking for Temporal Cascade Prediction arXiv cs.LG importance 50 / dev 60
- A Nesterov-Accelerated Byzantine-Robust Federated Learning arXiv cs.LG importance 55 / dev 70
- Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning arXiv cs.LG importance 45 / dev 60
- Adaptive Partitioning and Learning for Stochastic Control of Diffusion Processes arXiv cs.LG importance 45 / dev 55
- DuaDeep-SeqAffinity: Dual-Branch Deep Learning for Tri-Stream Sequence-Based Antibody--Antigen Affinity Prediction arXiv cs.LG importance 40 / dev 50
- Linearized subspace refinement framework to expose hidden accuracy in trained neural networks arXiv cs.LG importance 45 / dev 60
- MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration arXiv cs.LG importance 60 / dev 75
- Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators arXiv cs.LG importance 45 / dev 60
- Identification of Bivariate Causal Directionality Based on Anticipated Asymmetric Geometries arXiv cs.LG importance 35 / dev 45
- Sliding-Window Reordering with Overlap Averaging: A Simple Time-Domain Augmentation for Multivariate Forecasting arXiv cs.LG importance 40 / dev 55
- Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design arXiv cs.LG importance 65 / dev 70
- Towards Universal Tabular Embeddings: A Benchmark Across Data Tasks arXiv cs.LG importance 50 / dev 65
- RCProb: Probabilistic rule extraction from classification tree ensembles arXiv cs.LG importance 40 / dev 55
- Observation-Aligned Two-Stage Domain Decomposition for Physics-Informed Traffic State Estimation with Sparse Fixed Sensors arXiv cs.LG importance 45 / dev 60
- RW-TTT: Batched Serving for Request-Owned Test-Time Training State arXiv cs.LG importance 55 / dev 75
- Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning arXiv cs.LG importance 50 / dev 65
- Shortcuts in the Tail: Debiasing via Post-Hoc Spectral Compression of Fine-Tuning Updates arXiv cs.LG importance 50 / dev 65
- GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators arXiv cs.LG importance 45 / dev 55
- SPACR: Single-Pass Adaptive Training of Uncertainty-Aware Conformal Regressors arXiv cs.LG importance 45 / dev 60
- ThousandWorlds: A benchmark for climate emulation of potentially habitable exoplanets arXiv cs.LG importance 10 / dev 15
- Structured Inference with Large Language Gibbs arXiv cs.LG importance 40 / dev 60
- Real vs. Complex Spectral Bases for Neural Operators: The Role of Green's Function Alignment arXiv cs.LG importance 20 / dev 60
- Democratic ICAI: Debating Our Way to Steering Principles from Preferences arXiv cs.LG importance 45 / dev 75
- Target-Guided Selective Reweighting for Physics-Informed Neural Network Inverse Problems: A Transfer Learning Approach arXiv cs.LG importance 20 / dev 55
- Resample or Reroute? Recoverable Stopping Debt Without Identified Action Selection arXiv cs.LG importance 55 / dev 80
- ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling arXiv cs.LG importance 30 / dev 60
- Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling arXiv cs.LG importance 20 / dev 50
- Activation-Keyed Momentum: An Anisotropic Momentum Update via the Delta Rule arXiv cs.LG importance 25 / dev 65
- A Multidimensional Data-Driven Hybrid Transformer Framework for Non-invasive Continuous Blood Pressure Prediction arXiv cs.LG importance 15 / dev 45
- From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits arXiv cs.LG importance 20 / dev 45
- PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization arXiv cs.LG importance 25 / dev 65
- SimCast-S2S: A Computationally Efficient Diffusion Model for Subseasonal Precipitation Forecasting arXiv cs.LG importance 20 / dev 45
- Learning to Transfer Across Modes: Towards Unified Urban Mobility Forecasting arXiv cs.LG importance 20 / dev 55
- Hard-ReLU Gradient Descent Selects an Event-Free Sensitivity Limit arXiv cs.LG importance 25 / dev 75
- MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks arXiv cs.LG importance 35 / dev 75
- Modelpedia: A Catalog of Model Findings for the Meta-Science of AI arXiv cs.LG importance 55 / dev 85
- InKAN: B-Spline KANs via Truncated Power Form arXiv cs.LG importance 30 / dev 75
- Uncertainty Quantification in Machine Learning for Biosignal Applications -- A Review arXiv cs.LG importance 45 / dev 75
- Semiparametric Inference for Counterfactual Regression under Intervention-Driven Shift arXiv cs.LG importance 20 / dev 65
- Data-efficient Kernel Methods for Learning Hamiltonian Systems arXiv cs.LG importance 20 / dev 65
- Parameterized Hardness of Zonotope Containment and Neural Network Verification arXiv cs.LG importance 20 / dev 65
- Reliable Selection of Heterogeneous Treatment Effect Estimators arXiv cs.LG importance 20 / dev 55
- Active learning for data-driven reduced models of parametric differential systems with Bayesian operator inference arXiv cs.LG importance 25 / dev 65
- Non-Stationary Functional Bilevel Optimization arXiv cs.LG importance 20 / dev 65
- Deep networks learn to parse uniform-depth context-free languages from local statistics arXiv cs.LG importance 40 / dev 75
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories arXiv cs.LG importance 35 / dev 65
- Simplify to Amplify: Achieving Information-Theoretic Bounds with Fewer Steps in Spectral Community Detection arXiv cs.LG importance 20 / dev 65
- KernelFoundry: Hardware-aware evolutionary GPU kernel optimization arXiv cs.LG importance 55 / dev 85
- Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments arXiv cs.LG importance 20 / dev 55
- Learning to Concatenate Quantum Codes arXiv cs.LG importance 20 / dev 65
- Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification arXiv cs.LG importance 35 / dev 65
- A Real-Calibrated Synthetic-First Data Engine arXiv cs.LG importance 45 / dev 75
- The Timing Dependencies of Trust: Speed, Accuracy, and cBCI Neuro-Decoupling in Human-AI Teams arXiv cs.LG importance 35 / dev 55
- MidSurfNet: Learning Face Pairing for Mid-surface Abstraction of Thin-walled CAD Models arXiv cs.LG importance 20 / dev 55
- Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models arXiv cs.LG importance 45 / dev 85
- Explicit Interaction Architectures for Dynamical Learning: A Controlled Study of Structural Inductive Bias arXiv cs.LG importance 30 / dev 75
- Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations arXiv cs.LG importance 20 / dev 75
- Physics-Guided Robotic Radiation Source Localization along Arbitrary Measurement Paths in Unstructured Environments arXiv cs.LG importance 15 / dev 55
- What You See Is What You Get: Observation-Aligned Supervision for Chart-to-Code Generation arXiv cs.LG importance 50 / dev 85
- DrainSinkhorn: Safe Elimination for Batched Entropic Optimal Transport arXiv cs.LG importance 20 / dev 65
- Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO arXiv cs.LG importance 50 / dev 85
- Data Driven Equation Discovery for Phase-Ordering Dynamics : From Allen Cahn to the Ising Model arXiv cs.LG importance 25 / dev 65
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution arXiv cs.LG importance 65 / dev 85
- Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection arXiv cs.LG importance 40 / dev 75
- Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 arXiv cs.LG importance 65 / dev 85
- Improved Gradient Descent Lower Bounds Beyond Nesterov arXiv cs.LG importance 20 / dev 75
- Requirements After the First Edit: Mining Late Requirement Emergence and Rework in Real-World Coding-Agent Sessions arXiv cs.SE importance 60 / dev 85
- Large Language Models and Language Server Protocol: a match made in context arXiv cs.SE importance 55 / dev 85
- Compound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and Urgency arXiv cs.SE importance 55 / dev 85
- Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses arXiv cs.SE importance 50 / dev 80
- Refusing the Impossible: A Taxonomy and Benchmark for Code Hallucination in Large Language Models arXiv cs.SE importance 60 / dev 85
- TIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code Generation arXiv cs.SE importance 55 / dev 85
- Code Transformation Rule Synthesis using LLMs: Potential and Limits arXiv cs.SE importance 50 / dev 85
- No One Left Behind: Cross-Level Analysis for Sustainable Software Engineering arXiv cs.SE importance 40 / dev 65
- LabelMate: An LLM-Driven Framework for Refined Issue Report Labeling arXiv cs.SE importance 45 / dev 80
- ATIBA: Grounded Integrity and Quality Checking for Research Papers arXiv cs.SE importance 35 / dev 65
- The Illusion of Independent Quorums: Epistemic Fault Domains and Correlated Cognitive Failures in Agentic Quorums arXiv cs.SE importance 55 / dev 80
- Boundary-Mutation Testing for Pattern-Based Secret Detection: A Rule-Level Method and Cross-Scanner Evaluation arXiv cs.SE importance 45 / dev 80
- Virtual Testing of Automated Driving Systems through Credible Simulations arXiv cs.SE importance 35 / dev 65
- Quantisation of Abstract Data Types arXiv cs.SE importance 15 / dev 65
- CROCODIL: Cross-Model Code Editing with LLMs arXiv cs.SE importance 50 / dev 85
- SWIRL: Interactive Sensemaking of Tool-Generated Warnings through Customized Summaries arXiv cs.SE importance 45 / dev 80
- Extending Fill-In-the-Middle with Instructions for Steerable Code Completion arXiv cs.SE importance 55 / dev 85
- Detecting Multiple Semantic Concerns in Tangled Code Commits arXiv cs.SE importance 50 / dev 85
- A Longitudinal Study of Dependency Reclassifications in JavaScript Projects arXiv cs.SE importance 45 / dev 75
- LLM4Log: A Systematic Review of Large Language Model-based Log Analysis arXiv cs.SE importance 55 / dev 85
- Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub arXiv cs.SE importance 60 / dev 85
- A Physics-Informed Neuro-Fuzzy Framework for Quantum Error Attribution arXiv cs.SE importance 30 / dev 75
- The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the Browser arXiv cs.SE importance 50 / dev 85
- Surviving Code Reviews in the era of AI Lobsters importance 55 / dev 80
- .name Termination Lobsters importance 5 / dev 15
- Is AI ruining my brain? Lobsters importance 10 / dev 10
- The asteroid currently hitting frontend web development Lobsters importance 35 / dev 65
- The NX bit is not just about security Lobsters importance 25 / dev 55
- simple is not small Lobsters importance 20 / dev 45
- What are you doing this weekend? Lobsters importance 0 / dev 0
- smolts: a pedagogical IDE for a teaching language Lobsters importance 30 / dev 75
- Audacity 4.0.0 Released Lobsters importance 20 / dev 45
- Babashka 1.13.220 gets FFI Lobsters importance 25 / dev 75
- jank reimagines C++ errors and gets an official native package repo Lobsters importance 30 / dev 80
- Collider: An experimental Minecraft server Lobsters importance 15 / dev 35
- Revo Programming language Lobsters importance 20 / dev 75
- The new Go JSON API: twice as fast, or 1.5x slower? Lobsters importance 35 / dev 85
- Bug triaging help needed Lobsters importance 10 / dev 55
- Lua-async Lobsters importance 20 / dev 75
- A preview of the future Intel Architecture documentation Lobsters importance 25 / dev 55
- Imperial Colors Manifesto Lobsters importance 15 / dev 35
- jujutsu 0.45.0 Lobsters importance 25 / dev 85
- WTF is going on with R7RS Large? 2026 edition Lobsters importance 20 / dev 65
- FORCE_COLOR: Keep Your Command Line Colorful, Even when Piping - Retain Text Colors in Logs, Output Streams, and More Lobsters importance 20 / dev 75
- CERN transitioning industrial computers to Debian after being a longtime RHEL institution Lobsters importance 30 / dev 55
- The Incomplete Guide to Lazy Evaluation in Haskell (2015) Lobsters importance 20 / dev 75
- Inserting State Transitions in Postgres Lobsters importance 30 / dev 75
- AI-Powered Trading Strategies for Crypto Markets Dev.to AI importance 25 / dev 65
- Cómo una IA construye empresas sola Dev.to AI importance 30 / dev 65
- WordPress Is Using AI to Find Security Flaws Before Hackers Can Exploit Them Dev.to AI importance 45 / dev 75
- The 402 Wall — agents-only pixel billboard via x402 Dev.to AI importance 35 / dev 75
- Mastering Video Hooks with Logic Gap Dev.to AI importance 20 / dev 15
- Compress Image Online: Reduce Image Size with WebP & AVIF Dev.to AI importance 15 / dev 45
- Catch Tool Calls That Invent Missing Arguments Dev.to AI importance 55 / dev 75
- Secrets Stay Local Until the Hop Is Clean Dev.to AI importance 60 / dev 80
- The Tests Came With the Bug: Reviewing Agent PRs for Oracle Contamination Dev.to AI importance 55 / dev 70
- Workshop: Reject Incomplete Agent Plans Before Side Effects in 80 Minutes Dev.to AI importance 50 / dev 75
- Record the Hot Path, Then Extract One Seam Dev.to AI importance 55 / dev 80
- Add Screenshot Capability to Windsurf via MCP Dev.to AI importance 60 / dev 80
- The Hardest Part of a Proactive Assistant Is Knowing When Not to Speak Dev.to LLM importance 55 / dev 65
- Giving a Coding Agent an Org Chart Dev.to LLM importance 60 / dev 80
- Measuring the Multi-Agent Fork Tax Dev.to LLM importance 60 / dev 80
- I gave my drift monitor a denominator. The first thing it exposed was a hole in my own data collection Dev.to LLM importance 55 / dev 75
- Changes to LLM pricing: Baidu, StreamLake and Tencent Dev.to LLM importance 15 / dev 30
- Using LLMs for Crypto Market Analysis in 2026 Dev.to LLM importance 20 / dev 35
- LLMs: runbooks cortos para usar tools Dev.to LLM importance 60 / dev 80
- What to Delegate and What Not to Delegate to Local LLMs (with Benchmarks) Dev.to LLM importance 50 / dev 75
- Changes to LLM pricing: Baidu, NextBit and StreamLake Dev.to LLM importance 15 / dev 30
- Minimalistic OpenCode multi-agent configuration Dev.to LLM importance 55 / dev 80
- How AI Agents Are Changing the Way We Use the Internet Dev.to LLM importance 40 / dev 50
- How do you compare AI agents before committing to one? r/AI_Agents importance 30 / dev 40
- Kinda Funny... Coding Agent Follows Instructions Meant for the Agent it's Coding r/AI_Agents importance 35 / dev 60
- The 5 AI automations I'd build in any business this month (ranked by what they save) r/AI_Agents importance 35 / dev 45
- Is it just me or is AI terrible at building AI applications r/AI_Agents importance 40 / dev 65
- What made you trust a small tool enough to point it at your own files? r/AI_Agents importance 30 / dev 40
- Manager agent + worker agents in separate git worktrees: the orchestration patterns that survived contact with real overnight runs r/AI_Agents importance 60 / dev 80
- Did this happen to anyone? r/AI_Agents importance 40 / dev 45
- How are you learning Agentic AI right now- structured path or learn-as-you-go? r/AI_Agents importance 30 / dev 45
- I built my AI memory layer using itself in 10 days, here's what broke r/AI_Agents importance 60 / dev 80
- When an agent escapes its sandbox, where did the safeguards actually fail? r/AI_Agents importance 55 / dev 75
- Building a B2B SaaS for AI Agent Governance/Ops. Is "Human Cognitive Overload" the real bottleneck in production? Seeking builder feedback. r/AI_Agents importance 35 / dev 45
- AI Gateway Experiences r/AI_Agents importance 40 / dev 65
- Should RAG be agentic, or should the agent just decide where to retrieve from? r/AI_Agents importance 50 / dev 75
- Overwhelmed by all the AI options. What single subscription should I buy to build a basic MVP app? r/AI_Agents importance 25 / dev 30
- Any recommendations for building a real AI policy audit trail? r/AI_Agents importance 45 / dev 65
- Our real ceiling was tokens-per-minute, not latency: 2.2 turns/min for the whole product. When the queue saturates under that, do you drop the task or queue it? r/AI_Agents importance 60 / dev 80
- I Built an AI Chatbot That Was Useless. Then I Added a Teaching Layer. Now It's Actually Good. r/AI_Agents importance 40 / dev 60
- Ai snapchat bot. Can i break it? r/AI_Agents importance 20 / dev 25
- What is one AI agent workflow that businesses should automate before anything else? r/AI_Agents importance 35 / dev 45
- Looking for suggestions r/AI_Agents importance 20 / dev 25
- Question for people who manage / negotiate BPO or outsourcing contracts: what actually breaks with outcome-based pricing? r/AI_Agents importance 25 / dev 30
- I ACCIDENTALLY MADE A LOCAL RETRIEVAL TOOL WHICH I THINK IS WORKING! NEED HELP FOR IDEAS TO TEST IT r/AI_Agents importance 55 / dev 80
- 🚀 Two AI tools I’ve been working on r/AI_Agents importance 25 / dev 50
- My day today so far. r/OpenAI importance 5 / dev 10
- GPT-6-Astra's tax return underpays the government r/OpenAI importance 15 / dev 20
- Nice r/OpenAI importance 10 / dev 15
- He is a marketing Genius, knows how to keep the Hype going... r/OpenAI importance 10 / dev 15
- Turning myself into into random object pt 2 r/OpenAI importance 10 / dev 15
- I want the advanced model picker back r/OpenAI importance 15 / dev 25
- Downfall begins? r/OpenAI importance 15 / dev 20
- NeurIPS Sydney SOLD OUT in minutes [N] r/MachineLearning importance 25 / dev 35
- AAAI-27 desk rejection over incredibly minor abstract modifications [D] r/MachineLearning importance 20 / dev 30
- Grounding LLMs with JEPA-based world models trained in simulation — has this been tried? [D] r/MachineLearning importance 45 / dev 70
- Mol-JEPA - Multimodal molecular foundation model [R] r/MachineLearning importance 50 / dev 75
- How many repeated LLM queries are enough? Testing a pilot-based reliability protocol [R] r/MachineLearning importance 55 / dev 75
- August newsletter is out Simon Willison importance 30 / dev 45
- [AINews] Muse Spark 1.3 matches GPT-5.6-Sol, confirming Meta Superintelligence as the newest Frontier Lab, >90% discount for training Latent Space importance 50 / dev 75
- Inline Product Hunt importance 25 / dev 40
- Clockwork Product Hunt importance 25 / dev 45
- cmmnts Product Hunt importance 20 / dev 35
- Omarchy Product Hunt importance 30 / dev 50
- Chalked for Mac Product Hunt importance 25 / dev 45
- Snitch Product Hunt importance 20 / dev 35
- Omi Product Hunt importance 25 / dev 40
- Nex Product Hunt importance 30 / dev 50
- Study: Generative AI succumbs to conversational misinformed pressure and argument - news.arizona.edu Google News DeepSeek importance 40 / dev 60
- DeepSeek: How Has It Disrupted Global Open Source AI - AI Magazine Google News DeepSeek importance 45 / dev 70
- Only Codex Finished Austin Griffith's AI Security Race, DeepSeek Close Behind - Unchained - unchainedcrypto.com Google News DeepSeek importance 40 / dev 65
- Meta’s Flagship AI Model Makes Stunning Comeback: Cheaper Than DeepSeek, Chinese-American Top Executive Challenges Gemini Head-On - 36 Kr Google News DeepSeek importance 50 / dev 75
- How much longer can China afford cheap AI? - japantimes.co.jp Google News DeepSeek importance 35 / dev 50
- DeepSeek AI Price at $0.14 vs OpenAI’s $5 as Chinese Models Challenge $1.1 Trillion AI Boom - InfotechLead Google News DeepSeek importance 40 / dev 60
- Cracking 1.33 Trillion Daily Tokens: B.AI Powers the "AI Grid" with Full-Stack Infrastructure to Fuel the Agentic Era - GlobeNewswire Google News DeepSeek importance 50 / dev 75
- Open Models Now Handle 80-90% of Enterprise AI Tokens, Ollama's Jeffrey Morgan Says - finance.biggo.com Google News DeepSeek importance 50 / dev 75
- Saudi Arabia Built Its National AI Model on Top of China's MiniMax - Startup Fortune Google News DeepSeek importance 40 / dev 60
- From Gemini to GPT-6: 10 new AI models launched by Google, OpenAI and rivals - CNBC TV18 Google News DeepSeek importance 55 / dev 75
- Setting Grok Bot loose on procurement - X.ai Google News Grok/xAI importance 45 / dev 70
- Designing Grok Bot for a world of persistent agents - X.ai Google News Grok/xAI importance 50 / dev 75
- SpaceXAI expands Grok Bot to iPad as access expands to cheaper plans - 9to5Mac Google News Grok/xAI importance 40 / dev 60
- Musk’s xAI Fails to Block Minnesota’s AI ‘Nudification’ Law - Bloomberg Law News Google News Grok/xAI importance 45 / dev 70
- Tesla Optimus, Grok, Cybercab: Cutting Through the TSLA Hype - MarketWise Google News Grok/xAI importance 20 / dev 30
- Child sexual abuse survivor alleges Elon Musk’s AI chatbot used photos of her to generate new illegal images - theguardian.com Google News Grok/xAI importance 50 / dev 25
- xAI Is Offering Cheap Starlink to Angry Memphis Data Center Neighbors - yahoo.com Google News Grok/xAI importance 25 / dev 35
- Grok AI Is Coming to Tesla Robotaxi: What Riders Need to Know - BASENOR - Tesla Accessories Google News Grok/xAI importance 30 / dev 40
- ローカルLLMでベンチ豚(ベンチマーク豚)にならないためのHermes Agentの実践テク Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- ISO704を投げてAIの比喩をやめさせる Zenn LLM (いいね相当スコア: 0) importance 0 / dev 65
- llama.cpp の vkDestroyFence: Invalid device はメモリ不足ではない(Radeon 890M / Ry Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- RTX 5060 Ti 16GBでローカルLLM(Qwen3.8 27B)にもう一度挑んだ話 Zenn LLM (いいね相当スコア: 1) importance 0 / dev 75
- LLMでカタログPDFから構造化データを抽出する:プロンプト設計と後処理の実践 Zenn LLM (いいね相当スコア: 8) importance 0 / dev 80
- 3言語で字幕の遅延を測った — リズムは同じ、文字は違う Zenn LLM (いいね相当スコア: 0) importance 0 / dev 75
- 人間が書かない開発日誌の作り方 Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- Fable 5.1とGPT-6 Astraの定価が揃った。価格で選べなくなったのではなく、測らないと選べなくなった Zenn LLM (いいね相当スコア: 0) importance 0 / dev 75
- ContextPilot:32Kコンテキストで128Kを超える能動的文脈管理 Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- MCP Serverは何単位で分けるべきか? 統合と分割の設計軸を考える Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- Claude Code を7セッション並列で回すための設定を、全部置いておく Zenn LLM (いいね相当スコア: 0) importance 0 / dev 85
- Thinking(思考)をざっくり理解する Zenn LLM (いいね相当スコア: 0) importance 0 / dev 75
- Hugging Face事件から考える、AIエージェントが「引き返せなくなる」問題 Zenn LLM (いいね相当スコア: 0) importance 0 / dev 65
- Kimi K3の事後学習を技術報告書から読み解く Zenn LLM (いいね相当スコア: 10) importance 0 / dev 75
- Codex / Claude Codeの使用量を最適化する。Thinking・レビュー・コンテキスト運用見直しプロンプト Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- ローカルRAG環境を非エンジニアに届ける|ダブルクリックで入る「社内知恵袋(ShineosQA)」を作った話 Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- なぜ、喫茶店以外でスマホでChatGPT使う時は、質問したらすぐに画面を切らないといけないのか? Zenn LLM (いいね相当スコア: 2) importance 0 / dev 55
- 【Vol.25】SaaSがAIの中に入る日 — Claudeforceに学ぶ「AIから呼ばれる側」の設計 Zenn LLM (いいね相当スコア: 0) importance 0 / dev 80
- Metaの意地を見せたか Muse Glimmer 30Bの実力を見てみる Zenn LLM (いいね相当スコア: 1) importance 0 / dev 80
- RAGの精度を決める「チャンク分割」 — 食べすぎも食べなさすぎもダメ Zenn LLM (いいね相当スコア: 0) importance 0 / dev 85
- 過去の文章を元に、自然な「自分の文章」を書かせる doc-style-skill Zenn NLP (いいね相当スコア: 2) importance 0 / dev 65
- BigQuery MLの需要予測モデルでEC仕入れ量を最適化する実装手順 Zenn 機械学習 (いいね相当スコア: 0) importance 0 / dev 65
- MobileNetV2 を手書き NEON で速くする — 遅い層を個別に潰して ORT を抜く Zenn 機械学習 (いいね相当スコア: 0) importance 0 / dev 80
- AIに競艇予想モデルを作らせる。その数字を信じるのは人間の仕事です Zenn 機械学習 (いいね相当スコア: 0) importance 0 / dev 65
- 自動的に格納したファイルを適切なフォルダに格納するにはどう判定していくか考えてみる Zenn 機械学習 (いいね相当スコア: 1) importance 0 / dev 70
- GPT-2-likeからQwen2-likeへの実験:第2回 GELUをSwiGLUに変える Zenn 機械学習 (いいね相当スコア: 3) importance 0 / dev 75
- [レビュー] GoogleのTimesFM-3をリリース直後に自分で検証しました:ゼロショットは本物、目玉機能はまだ Zenn 機械学習 (いいね相当スコア: 3) importance 0 / dev 70
- Kaggle AI Agent Securityコンペ振り返り ー345th Place Solution Zenn 機械学習 (いいね相当スコア: 6) importance 0 / dev 70
- なぜ、OpenAIのGPT-6 Astraの広告映像は滑稽なのか? Qiita LLM (いいね相当スコア: 0) importance 0 / dev 30
- リアルタイム音声AIが同じ発話に2回答するのを防ぐ:STT確定ログをPythonで実装する Qiita LLM (いいね相当スコア: 0) importance 0 / dev 75
- LLMは魔法じゃない — 本番で構造化をやって踏んだ4つの穴(既存のテストでは拾えない) Qiita LLM (いいね相当スコア: 0) importance 0 / dev 75
- Claude Code v2.1.260の重大セキュリティ修正と/diffパネル追加を解説 Qiita LLM (いいね相当スコア: 1) importance 0 / dev 85
- 因果推論 Day 22/全30回 効果の異質性、平均の裏で誰に効いて誰に効かないか Qiita 機械学習 (いいね相当スコア: 0) importance 0 / dev 70
- 因果推論 Day 21/全30回 ITTとLATE、割付は守られないという現実 Qiita 機械学習 (いいね相当スコア: 0) importance 0 / dev 70
- スポーツ分析で活用されるAIモデル入門:データからパフォーマンスを読み解く方法 Qiita 機械学習 (いいね相当スコア: 1) importance 0 / dev 60
- 【技術解説】【完全ガイド】Pythonで株の自動売買システムを構築する方法 Qiita 機械学習 (いいね相当スコア: 0) importance 0 / dev 65
- LLMは相対性理論を独立して発見できない、という意見どうよ?相棒。 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 40
- Repository Explorer – 大胆にリエンジニアリングする note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 60
- AIと非難なき責任 / AIの無過失損害と偶発的関係 雑感 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 20
- 法に従うAIと法を破る自由 / プリンシパル・スイッチの構想 雑感 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 15
- ローカルLLM用マシンの選び方――メモリ容量・対応環境・速度で比較する note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- AIと選挙 / 自律性のパラドックス / 真偽を単位とする選挙偽情報規制の外側 雑感 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 10
- 第三話「慎重すぎるクロエちゃん」 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 20
- 減らせと自分で書いた日から、AIへの指示書は60%増えた note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 70
- S3-10 AIごとに、得意な仕事を分ければいい note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- GPT-6を待つ前に。AIのコード品質を守っていたのは「仕組み」だった note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 80
- LLMルーター運用ログ:2026年9月。変更「0行」という確信を得るための定点観測 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- 現状LLMにはアブダクションはできない、けれども。 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 50
- 最新AI....なぜ、必ず1回止まるのか? note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 30
- AIやらかし30例と全直し方フル版 送信前1行検算で防ぐ業務ミスの実践カタログ note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 70
- LaughingCloudとは note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 40
- 身体にも「ChatGPT」のようなAIが生まれる?Apple Watchと「身体のLLM」の未来 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 45
- 「どのAIモデルを使うか」を代行する会社に、Stripeが約1兆円を払った理由 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- 買った本|ロールプレイの本質を知る note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 25
- RIFA-FLASH-1.7B 総合ベンチマークレポート(全52問) note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- RIFA-FLASH-1.7B Phase 5 インジェクション後編レポート note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 80
- AI時代のIT世界の現実 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- 鳴った(Opus4.6) note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 25
- AIは人間の言葉を離れて考えはじめるのか ― 潜在推論、Recurrent Depth、COCONUTから「機械固有の認知」へ note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 60
- だいたい全部、ベクトルで考える note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 40
- 🔊音声あり(日&英):AI音声合成の新境地!「演技指導」に従って声のニュアンスを自在に変える最新TTS技術を解説【論文解説】 note LLM (いいね相当スコア: 取得失敗) importance 0 / dev 75
- Java誕生から現代に至る歴史をゴスリン氏ら関係者が語る公式ドキュメンタリー「The Java Story」、YouTubeで公開中 Publickey importance 30 / dev 50
- Webブラウザ上でターミナルのシミュレータを実行、GitやCLI、Vimなどを無料で学べる「WebTerm Learn」が公開 Publickey importance 40 / dev 65
- AWS、新たな太平洋海底ケーブル「Sta'O'Nuk」発表。420Tbps、2029年に稼働へ。集中リスク排除のため新たな陸揚げ拠点を採用 Publickey importance 40 / dev 50
- iOSDC Japan 2026にLINEヤフーのエンジニア7名が登壇します(9/11〜9/13) LY Corp Tech Blog importance 15 / dev 35