AI News Digest 2026-08-27
直近2日間のAI関連ニュースから、番組で扱った記事と収集した全記事を一覧にしたものです。
台本で使った記事
特集
開発者コーナー
中堅コーナー
ハーネスコーナー
速報コーナー
参考記事一覧
参考記事一覧を表示(925件)
- GLM-5.3-Flash: How Z.ai Built a 320B MoE That Runs at 1/10th the Cost of Its Predecessor Dev.to LLM / imp 80
- [Megathread] Qwen3.8-Flash-Next - Release Day r/LocalLLaMA / imp 85
- OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks The Decoder / imp 0
- Bill Gates warns AI is more dangerous than the tech industry will admit The Decoder / imp 35
- Russia used ChatGPT to run a covert influence campaign pushing pro-Kremlin narratives across the West The Decoder / imp 25
- IBM drops open-weight Granite 4.2 family with built-in agentic capabilities under Apache 2.0 The Decoder / imp 75
- Google’s new AI transcription edits out your 'ums' and 'ahs' The Verge (AI) / imp 35
- Employee revolt and failing agents forced Meta to scrap its AI layoff plan The Decoder / imp 25
- Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory r/LocalLLaMA / imp 25
- DeepSeek Looks to Raise $7 Billion as Revenues Jump Tenfold - PYMNTS.com Google News DeepSeek / imp 55
- DeepSeek targets 2027 listing as pre-IPO funding nears close: sources - South China Morning Post Google News DeepSeek / imp 50
- DeepSeek, ChatGPT and Claude: How Chinese hackers are using AI in cyberattacks - Firstpost Google News DeepSeek / imp 20
- China's richest man, Zhong Shanshan, invests $350 million in DeepSeek AI - KuCoin Google News DeepSeek / imp 45
- Open Source DeepSeek Harness Launches with an MIT License - Geeky Gadgets Google News DeepSeek / imp 15
- Musk admits Grok lags rivals, elevates Anthropic as AI leader - CHOSUNBIZ - Chosunbiz Google News Grok/xAI / imp 30
- SpaceXAI will deploy standalone Nvidia Vera CPUs for Grok's agentic workloads — will use optimized Vera Rubin NVL72 in space with Starmind satellite - Tom's Hardware Google News Grok/xAI / imp 40
- Elon Musk Says Grok Is Taking Over Starlink Support: 15,000+ Calls a Day, 3,000+ Orders a Week - Benzinga Google News Grok/xAI / imp 20
- JPMorgan Is ‘Increasingly Positive’ On SpaceX’s Grok-Cursor Push – Sees 75% Potential Upside In SPCX Stock - Stocktwits Google News Grok/xAI / imp 25
- Scalable Capital Opens Its Brokerage Accounts to AI Agents like Claude, Grok and ChatGPT - trendingtopics.eu Google News Grok/xAI / imp 40
- EVE Online: The Move to Python 3 Begins! Simon Willison / imp 50
- Two humanoid robots broke Usain Bolt's world record this week. That's not even the most interesting thing that happened in AI. r/artificial / imp 35
- v2.1.246 Claude Code / imp 45
- Tailcat Hacker News / imp 35
- YouTube Format IDs Hacker News / imp 15
- v1.18.23 OpenCode / imp 40
- Bringing ChatGPT for Teachers to more U.S. school districts OpenAI News / imp 30
- Learning never stops: How AI makes learning continuous OpenAI News / imp 25
- How loveholidays is making everyone a builder with Codex OpenAI News / imp 40
- The full stack behind abundant intelligence OpenAI News / imp 60
- Introducing the Admin plugin for ChatGPT Work and Codex OpenAI News / imp 35
- 5 ways to upgrade your home decor with Google Search Google AI Blog / imp 5
- GlucoFM: Foundation model for continuous glucose monitoring Google Research Blog / imp 45
- AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR Google Research Blog / imp 30
- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers Hugging Face Blog / imp 50
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original Hugging Face Blog / imp 55
- Wire It, Run It, Deploy It: AI Workflows in Gradio Hugging Face Blog / imp 45
- How to evaluate LLMs before production GitHub Blog / imp 55
- How Much Code Do Developers Really Let Agents Write? JetBrains Blog / imp 50
- AI Agents in DataGrip JetBrains Blog / imp 55
- Compose Multiplatform 1.12.0 Released JetBrains Blog / imp 60
- How Ubuntu Is Using Rust to Rebuild Core System Tools JetBrains Blog / imp 50
- OpenTelemetry Comes to IntelliJ IDEA, GoLand, PyCharm, and WebStorm JetBrains Blog / imp 45
- Ideas Worth a Longer Conversation: The JetBrains Research Podcast JetBrains Blog / imp 35
- Moving from Minimus to Docker Hardened Images Docker Blog / imp 40
- Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding NVIDIA Developer Blog / imp 55
- Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo NVIDIA Developer Blog / imp 50
- NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI NVIDIA Developer Blog / imp 60
- Raised on AI MIT Technology Review (AI) / imp 15
- AI models flub these intelligence tests. Can you fare any better? MIT Technology Review (AI) / imp 30
- AI won’t replace radiologists, but it will dramatically change their jobs Ars Technica (AI) / imp 25
- Radar makes podcasts searchable — and usable by AI agents TechCrunch (AI) / imp 55
- Ex-Meta scientists want to bring visual AI to the factory floor TechCrunch (AI) / imp 35
- Robot brain builders are pushing out of their GPT-2 era TechCrunch (AI) / imp 25
- QueryStory wants you to believe what AI is telling you TechCrunch (AI) / imp 30
- Arga Labs is building a better way to train enterprise AI agents TechCrunch (AI) / imp 45
- Hearing tech startup Legato emerges from stealth with $12M and a peek at its AI hearing glasses TechCrunch (AI) / imp 25
- Runable hits $21M to bet AI agents can go from building businesses to growing them TechCrunch (AI) / imp 40
- India's Ringg gets backing from Peak XV as it pushes voice AI past the phone call TechCrunch (AI) / imp 25
- Robotics startup Generalist reaches $3B valuation, sources say TechCrunch (AI) / imp 35
- OpenAI loses a top data center exec as stream of high-profile departures continues TechCrunch (AI) / imp 25
- Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding TechCrunch (AI) / imp 50
- Claude Cowork finally remembers what you told the app in chat TechCrunch (AI) / imp 65
- Gamma acquires Accel-backed design startup Lica TechCrunch (AI) / imp 30
- Accel-backed Keenable is indexing the web for AI agents TechCrunch (AI) / imp 50
- 'The world seems to be ready': An interview with OpenAI head of product Thibault Sottiaux TechCrunch (AI) / imp 35
- Situational Awareness, star AI hedge fund that nearly imploded, now being probed by the SEC TechCrunch (AI) / imp 20
- OpenAI subpoenaed by Alabama AG over Hugging Face hack The Verge (AI) / imp 50
- Presentation: Can Claude Fix Itself? Using LLMs for Incident Response InfoQ (AI/ML/Data Eng) / imp 55
- Article: Beyond Offset Lag: Computing Time in Queue for Apache Hudi Data Lake Pipelines at Petabyte Scale InfoQ (AI/ML/Data Eng) / imp 40
- Diagrid Catalyst 2.0 Adds Durable and Verifiable Execution for AI Agents InfoQ (AI/ML/Data Eng) / imp 50
- Cursor Releases Origin as an Agent-Native Alternative to GitHub InfoQ (AI/ML/Data Eng) / imp 0
- Beyond Embedded: How DuckDB v2.0 Shifts Architecture toward Distributed Network Capabilities InfoQ (AI/ML/Data Eng) / imp 60
- New Platform Peers Inside AI’s Black Box IEEE Spectrum (AI) / imp 40
- AI Companion Robots Are Closing the Human Connection in Modern Homes IEEE Spectrum (AI) / imp 20
- Sam Altman says OpenAI will have AGI by the end of 2026 if you accept his definition The Decoder / imp 65
- Pro-Kremlin deepfakes put surrender rhetoric in the mouths of Ukrainian lawmakers The Decoder / imp 15
- Chinese Moonshot AI negotiates hosting deals with Microsoft, Amazon, and Google The Decoder / imp 45
- Anthropic sees a market opportunity of more than $30 trillion ahead of its IPO The Decoder / imp 50
- Last Week in AI #342 - Last 3 Months in AI Last Week in AI / imp 25
- RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation arXiv cs.AI / imp 45
- ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence arXiv cs.AI / imp 50
- LLM Agents Perform Controlled Experiments Using Simulation Models arXiv cs.AI / imp 55
- A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts arXiv cs.AI / imp 45
- TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery arXiv cs.AI / imp 55
- Function-Level Execution Feedback for Code Preference Optimization arXiv cs.AI / imp 50
- Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes arXiv cs.AI / imp 30
- How much of a measured AI preference is the model, and how much is the instrument? arXiv cs.AI / imp 35
- AI Agents Push Humans Out of the Loop arXiv cs.AI / imp 40
- FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare arXiv cs.AI / imp 45
- Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value arXiv cs.AI / imp 55
- MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models arXiv cs.AI / imp 50
- Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering arXiv cs.AI / imp 45
- Automata from Agent Traces: Failure and Next-Step Prediction arXiv cs.AI / imp 55
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment arXiv cs.AI / imp 60
- Do LLMs Understand Limit Order Book Dynamics? arXiv cs.AI / imp 40
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace arXiv cs.AI / imp 55
- Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware arXiv cs.AI / imp 50
- Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search arXiv cs.AI / imp 50
- A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification arXiv cs.AI / imp 45
- Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention arXiv cs.AI / imp 55
- SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models arXiv cs.AI / imp 50
- Exploit More, Explore Smarter for Budget-Constrained Agentic Search arXiv cs.AI / imp 55
- In-Context Inpainting for Time Series Forecasting arXiv cs.AI / imp 40
- Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications arXiv cs.AI / imp 65
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors arXiv cs.AI / imp 65
- AI Finds A Way arXiv cs.AI / imp 45
- Provenance Guided Incremental Learning Under Evolving Concept Definitions arXiv cs.AI / imp 40
- BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification arXiv cs.AI / imp 45
- Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems arXiv cs.AI / imp 60
- Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment arXiv cs.AI / imp 45
- PROOF-Gen: From Optimized Data to Better Distillation arXiv cs.AI / imp 55
- MARS: Multi-Specialist LLM Relay System for Competitive Programming arXiv cs.AI / imp 45
- Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining arXiv cs.AI / imp 55
- Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors arXiv cs.AI / imp 45
- More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight arXiv cs.AI / imp 65
- Recursive Agentic Reasoning arXiv cs.AI / imp 55
- More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving arXiv cs.AI / imp 55
- Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design arXiv cs.AI / imp 45
- When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs arXiv cs.AI / imp 45
- Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling arXiv cs.AI / imp 35
- Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning arXiv cs.AI / imp 45
- Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction arXiv cs.AI / imp 45
- Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing arXiv cs.AI / imp 40
- Reflection with Action-Induced Visual Differences for Desktop GUI Agents arXiv cs.AI / imp 55
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding arXiv cs.AI / imp 55
- Relative Time Intervals Representation for Word-level Timestamping with Masked Training arXiv cs.AI / imp 45
- Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment arXiv cs.AI / imp 55
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems arXiv cs.AI / imp 65
- Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression arXiv cs.AI / imp 55
- AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval arXiv cs.AI / imp 55
- EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals arXiv cs.AI / imp 45
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments arXiv cs.AI / imp 55
- ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation arXiv cs.AI / imp 55
- Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing arXiv cs.AI / imp 50
- AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL arXiv cs.AI / imp 55
- Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping arXiv cs.AI / imp 60
- OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses arXiv cs.AI / imp 45
- Task-Adaptive Rubrics for GUI Reward Modeling arXiv cs.AI / imp 55
- Paritok-4B: Intent-Conditioned Context Compression for Coding Agents arXiv cs.AI / imp 60
- Preference Data Selection for Mitigating the Alignment Tax in Large Language Models arXiv cs.AI / imp 60
- MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG arXiv cs.AI / imp 60
- Constraint-Guided Enterprise Data Mapping with Large Language Models arXiv cs.AI / imp 50
- Evaluating Multiple LLM Generations with Validated Task Coverage arXiv cs.AI / imp 50
- TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models arXiv cs.AI / imp 60
- STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation arXiv cs.AI / imp 50
- SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction arXiv cs.AI / imp 55
- Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks arXiv cs.AI / imp 45
- Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing arXiv cs.AI / imp 40
- Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding arXiv cs.AI / imp 45
- Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling arXiv cs.AI / imp 40
- RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards arXiv cs.AI / imp 65
- ReproAgent: Contract-Guided Paper-to-Code Reproduction arXiv cs.AI / imp 60
- VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models arXiv cs.AI / imp 60
- OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning arXiv cs.AI / imp 60
- Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight arXiv cs.AI / imp 55
- Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control arXiv cs.AI / imp 50
- SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception arXiv cs.AI / imp 45
- Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning arXiv cs.AI / imp 60
- The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents arXiv cs.AI / imp 60
- Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems arXiv cs.AI / imp 60
- From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use arXiv cs.AI / imp 65
- Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs arXiv cs.AI / imp 40
- ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping arXiv cs.AI / imp 55
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation arXiv cs.AI / imp 50
- Partial Identification under Causal Orders by Linear Programming arXiv cs.AI / imp 40
- A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads arXiv cs.AI / imp 40
- Mahalanobis-Based Multi-Head Attention for Complex State Propagation arXiv cs.AI / imp 45
- HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning arXiv cs.AI / imp 40
- Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling arXiv cs.AI / imp 40
- Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites arXiv cs.AI / imp 40
- PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents arXiv cs.AI / imp 60
- Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models arXiv cs.AI / imp 60
- Discovering Adaptive Transmission Programs for Collective Innovation arXiv cs.AI / imp 40
- When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows arXiv cs.AI / imp 60
- EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents arXiv cs.AI / imp 60
- Joint Optimization of Tool Creation and Use for Large Language Model Agents arXiv cs.AI / imp 65
- PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos arXiv cs.AI / imp 50
- Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms arXiv cs.AI / imp 50
- Causal Modelling of Support Interventions for Student Competency Assessment arXiv cs.AI / imp 40
- Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning arXiv cs.AI / imp 60
- The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models arXiv cs.AI / imp 60
- Confident at the moment of action: belief miscalibration in LLM play under hidden information arXiv cs.AI / imp 50
- Lifted Model Construction under Approximate Commutativity arXiv cs.AI / imp 40
- Meta$^n$: Recursive Self-Improvement through Emergent Depth arXiv cs.AI / imp 65
- RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons arXiv cs.AI / imp 50
- Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav arXiv cs.AI / imp 60
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing arXiv cs.AI / imp 65
- Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought arXiv cs.AI / imp 55
- CAFE: Self-Improving Search Agents Need Co-Evolving Feedback arXiv cs.AI / imp 60
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments arXiv cs.AI / imp 65
- Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core arXiv cs.AI / imp 45
- Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA arXiv cs.AI / imp 50
- A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments arXiv cs.AI / imp 40
- FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs arXiv cs.AI / imp 50
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL arXiv cs.AI / imp 60
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses arXiv cs.AI / imp 65
- Progressively Learning Heterogeneous Skills in a Unified Latent Space arXiv cs.AI / imp 55
- A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision arXiv cs.AI / imp 40
- Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers arXiv cs.AI / imp 55
- REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring arXiv cs.AI / imp 60
- Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal arXiv cs.AI / imp 50
- Identifying Latent Declarative Representations of Code for Assisting Repository Migration arXiv cs.AI / imp 45
- When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs arXiv cs.AI / imp 55
- Macro-Operator Generation and Predicate Selection for TAMP Operator Learning arXiv cs.AI / imp 30
- ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents arXiv cs.AI / imp 60
- Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail arXiv cs.AI / imp 65
- Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling arXiv cs.AI / imp 50
- Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap arXiv cs.AI / imp 50
- From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers arXiv cs.AI / imp 45
- Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model arXiv cs.AI / imp 50
- The Limits of Automatic Evaluation of Creativity in Large Language Models arXiv cs.AI / imp 40
- Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it arXiv cs.AI / imp 45
- EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$ arXiv cs.AI / imp 40
- TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers arXiv cs.AI / imp 75
- What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development arXiv cs.AI / imp 35
- Disentangled Skill Representations for Predictive Human Modeling arXiv cs.AI / imp 45
- When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk arXiv cs.AI / imp 40
- EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis arXiv cs.AI / imp 40
- Restoring Without Forgetting: Continual Learning Across Image Degradations arXiv cs.AI / imp 45
- LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology arXiv cs.AI / imp 50
- Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders arXiv cs.AI / imp 45
- Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring arXiv cs.AI / imp 55
- Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB arXiv cs.AI / imp 40
- Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation arXiv cs.AI / imp 35
- Infant Care Video Dataset for Classification of Interventions Using Transformers arXiv cs.AI / imp 35
- Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization arXiv cs.AI / imp 60
- ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork arXiv cs.AI / imp 55
- Automated Synthesis of Cloud Emulators arXiv cs.AI / imp 55
- Coronavirus Optimization Algorithm: A Success-History Adaptive Evolutionary Framework with Archive-Assisted Search and Stagnation Recovery for Global Optimization arXiv cs.AI / imp 30
- Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2) arXiv cs.AI / imp 70
- Revelation Control arXiv cs.AI / imp 40
- A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization arXiv cs.AI / imp 50
- A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource arXiv cs.AI / imp 35
- Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory arXiv cs.AI / imp 45
- Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs arXiv cs.AI / imp 70
- RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding arXiv cs.AI / imp 35
- QML for Quantum Sensing under Measurement-Induced Information Loss arXiv cs.AI / imp 40
- Luce: Relightable Gaussians for 3D Asset Generation arXiv cs.AI / imp 45
- STAIN-FL: Stealthy Targeted Attack Injection with Contextual Triggers in Federated Learning arXiv cs.AI / imp 55
- The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses arXiv cs.AI / imp 75
- NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution arXiv cs.AI / imp 60
- Evaluating Language Models on Cross-Language Code Functional Equivalence arXiv cs.AI / imp 50
- RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation arXiv cs.AI / imp 70
- The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem arXiv cs.AI / imp 60
- Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation arXiv cs.AI / imp 70
- SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding arXiv cs.AI / imp 50
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents arXiv cs.AI / imp 70
- IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views arXiv cs.AI / imp 50
- What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions arXiv cs.AI / imp 75
- ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning arXiv cs.AI / imp 50
- Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings arXiv cs.AI / imp 55
- Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models arXiv cs.AI / imp 50
- Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding arXiv cs.AI / imp 55
- VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference arXiv cs.AI / imp 55
- Mechanistic Circuit Identification for Controllable Data Generation arXiv cs.AI / imp 45
- ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal arXiv cs.AI / imp 40
- When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs arXiv cs.AI / imp 45
- PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding arXiv cs.AI / imp 55
- Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents arXiv cs.AI / imp 70
- MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes arXiv cs.AI / imp 40
- Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection arXiv cs.AI / imp 55
- PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control arXiv cs.AI / imp 55
- TransPhy: Visual In-Context Learning for Physically Grounded Image Editing arXiv cs.AI / imp 45
- Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking arXiv cs.AI / imp 45
- From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender arXiv cs.AI / imp 60
- PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment arXiv cs.AI / imp 45
- Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection arXiv cs.AI / imp 45
- LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes arXiv cs.AI / imp 55
- Preference Optimization for Non-Verbal Vocalization Synthesis arXiv cs.AI / imp 45
- Tlow: Flow-based Item Tokenizer for Recommendation arXiv cs.AI / imp 50
- 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection arXiv cs.AI / imp 65
- Contrastive Branch Policy Optimization arXiv cs.AI / imp 60
- SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling arXiv cs.AI / imp 45
- Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning arXiv cs.AI / imp 45
- Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis arXiv cs.AI / imp 50
- FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision arXiv cs.AI / imp 65
- Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs arXiv cs.AI / imp 60
- Markerless Pose Estimation for Resistance Training Technique Assessment arXiv cs.AI / imp 40
- Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning arXiv cs.AI / imp 40
- Multilevel Fair Allocation under Additive Preferences arXiv cs.AI / imp 35
- Evaluating Deep Multivariate Imputation Models on Wearable Device Data arXiv cs.AI / imp 45
- Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning arXiv cs.AI / imp 55
- When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study arXiv cs.AI / imp 65
- Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR arXiv cs.AI / imp 45
- LumiXAI: A Modular Full-Stack Framework for Feature Attribution arXiv cs.AI / imp 55
- FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment arXiv cs.AI / imp 60
- StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment arXiv cs.AI / imp 55
- COCI: Conference Organisers and Content Identifier arXiv cs.AI / imp 40
- Across the Loss Landscape with Progressive Growth arXiv cs.AI / imp 40
- $\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions arXiv cs.AI / imp 55
- Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks arXiv cs.AI / imp 50
- A Literate Programming Environment for Human and Machine Agents arXiv cs.AI / imp 60
- Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems arXiv cs.AI / imp 65
- Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration arXiv cs.AI / imp 60
- On-policy Distillation with Verifiable Reward arXiv cs.AI / imp 65
- Constrained Hyperparameter Optimization for Streaming Data arXiv cs.AI / imp 45
- Deep Learning Super Resolution for Satellite Cloud Mask Downscaling arXiv cs.AI / imp 40
- Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity arXiv cs.AI / imp 50
- Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets arXiv cs.AI / imp 50
- Method, Mind, and Morality: How People Make Sense of Artificial Intelligence arXiv cs.AI / imp 40
- The RAT: A Unified Bayesian Model for RAG Evaluation arXiv cs.AI / imp 40
- Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations arXiv cs.AI / imp 20
- Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks arXiv cs.AI / imp 15
- Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy arXiv cs.AI / imp 30
- Automatic Model Card Generation Using an LLM arXiv cs.AI / imp 40
- Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows arXiv cs.AI / imp 50
- LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training arXiv cs.AI / imp 40
- Fuzzy Segmentations of a String arXiv cs.AI / imp 15
- Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives arXiv cs.AI / imp 30
- Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs arXiv cs.AI / imp 35
- LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis arXiv cs.AI / imp 45
- Efficient LLM Collaboration via Planning arXiv cs.AI / imp 65
- Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light arXiv cs.AI / imp 40
- Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning arXiv cs.AI / imp 35
- UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models arXiv cs.AI / imp 40
- ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering arXiv cs.AI / imp 55
- Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge arXiv cs.AI / imp 40
- Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models arXiv cs.AI / imp 50
- CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support arXiv cs.AI / imp 45
- PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools arXiv cs.AI / imp 70
- Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints arXiv cs.AI / imp 35
- ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams arXiv cs.AI / imp 35
- Housing Potential Common Data Model and City Digital Twin arXiv cs.AI / imp 25
- Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust arXiv cs.AI / imp 45
- EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design arXiv cs.AI / imp 65
- Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control arXiv cs.AI / imp 60
- Atomic Units of X: The Compression Layer of Intelligence arXiv cs.AI / imp 50
- SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents arXiv cs.AI / imp 70
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models arXiv cs.AI / imp 55
- The Dynamics of Intelligence Explosions arXiv cs.AI / imp 65
- Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument arXiv cs.AI / imp 50
- Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies arXiv cs.AI / imp 35
- ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting arXiv cs.AI / imp 30
- What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies arXiv cs.AI / imp 35
- SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing arXiv cs.AI / imp 35
- GenCoord: Skill-Path Commitments under Private Information arXiv cs.AI / imp 50
- Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models arXiv cs.AI / imp 40
- CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents arXiv cs.AI / imp 55
- SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems arXiv cs.AI / imp 35
- MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks arXiv cs.AI / imp 65
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work arXiv cs.AI / imp 0
- MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction arXiv cs.AI / imp 50
- Screening Autism Spectrum Disorder in children using Deep Learning Approach : Evaluating the classification model of YOLOv26s by comparing with other models arXiv cs.AI / imp 30
- HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA arXiv cs.AI / imp 45
- Intrinsic PAPR: Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering arXiv cs.AI / imp 25
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control arXiv cs.AI / imp 40
- Generative AI for Validating Physics Laws arXiv cs.AI / imp 40
- Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review arXiv cs.AI / imp 55
- Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric arXiv cs.AI / imp 40
- Quasar: A Programming Language Specialized for LLM Code Actions arXiv cs.AI / imp 65
- From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs arXiv cs.AI / imp 45
- A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs arXiv cs.AI / imp 50
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities arXiv cs.AI / imp 40
- NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs arXiv cs.AI / imp 55
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification arXiv cs.AI / imp 50
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification arXiv cs.AI / imp 30
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships arXiv cs.AI / imp 40
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models arXiv cs.AI / imp 40
- Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics? arXiv cs.AI / imp 35
- Monotone and Separable Set Functions: Characterizations and Neural Models arXiv cs.AI / imp 20
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution arXiv cs.AI / imp 40
- Robust Motion Generation using Part-level Reliable Data from Videos arXiv cs.AI / imp 35
- Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library arXiv cs.AI / imp 40
- Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations arXiv cs.AI / imp 55
- Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes arXiv cs.AI / imp 45
- Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model arXiv cs.AI / imp 30
- TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning arXiv cs.AI / imp 40
- Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility arXiv cs.AI / imp 45
- Ad Insertion in LLM-Generated Responses arXiv cs.AI / imp 35
- Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging arXiv cs.AI / imp 50
- ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents arXiv cs.AI / imp 45
- PatientHub: A Unified Framework for Patient Simulation arXiv cs.AI / imp 45
- You Can Learn Tokenization End-to-End with Reinforcement Learning arXiv cs.AI / imp 55
- VLANeXt: Recipes for Building Strong VLA Models arXiv cs.AI / imp 60
- ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents arXiv cs.AI / imp 60
- EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training arXiv cs.AI / imp 30
- ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models arXiv cs.AI / imp 60
- msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models arXiv cs.AI / imp 40
- Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models arXiv cs.AI / imp 45
- Beyond OAuth: Task-Scoped Authorization for AI Agents via Natural Language Slices arXiv cs.AI / imp 70
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification arXiv cs.AI / imp 40
- Ollivier-Ricci Curvature of Riemannian Manifolds and Directed Graphs with Applications to Graph Neural Networks arXiv cs.AI / imp 20
- Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts arXiv cs.AI / imp 40
- RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction arXiv cs.AI / imp 35
- Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters arXiv cs.AI / imp 50
- Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval arXiv cs.AI / imp 65
- Outlier-Robust Diffusion Solvers for Inverse Problems arXiv cs.AI / imp 35
- CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving arXiv cs.AI / imp 60
- ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching arXiv cs.AI / imp 45
- Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation arXiv cs.AI / imp 55
- Skill-Conditioned Gated Self-Distillation for LLM Reasoning arXiv cs.AI / imp 50
- A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink arXiv cs.AI / imp 45
- SaliMory: Orchestrating Cognitive Memory for Conversational Agents arXiv cs.AI / imp 70
- When Can One Neuron Fix Repetition Loops in LLMs? arXiv cs.AI / imp 50
- TW-LegalBench: Measuring Taiwanese Legal Understanding arXiv cs.AI / imp 35
- RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation arXiv cs.AI / imp 55
- Co-occurring Associated REtained concepts in Diffusion Unlearning arXiv cs.AI / imp 40
- Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop arXiv cs.AI / imp 70
- A Unified Algebraic Framework for Classification Performance Evaluation arXiv cs.AI / imp 40
- Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution arXiv cs.AI / imp 75
- The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle arXiv cs.AI / imp 40
- Discrete Diffusion Models: A Unified Framework from Tokenization to Generation arXiv cs.AI / imp 50
- EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology arXiv cs.AI / imp 30
- GraphVid: Interactive Graph-Controllable Video Generation arXiv cs.AI / imp 35
- LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding arXiv cs.AI / imp 60
- MOSAIC: Masked Outsourcing of Secure AI Computations arXiv cs.AI / imp 45
- TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction arXiv cs.AI / imp 50
- Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset arXiv cs.AI / imp 25
- AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization arXiv cs.AI / imp 35
- Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol arXiv cs.AI / imp 55
- ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning arXiv cs.AI / imp 45
- MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation arXiv cs.AI / imp 60
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation arXiv cs.AI / imp 50
- Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas arXiv cs.AI / imp 30
- Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay arXiv cs.AI / imp 65
- ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations arXiv cs.AI / imp 35
- Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation arXiv cs.AI / imp 40
- Scaling Muon for Diffusion Transformers arXiv cs.AI / imp 45
- CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents arXiv cs.AI / imp 55
- Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds arXiv cs.AI / imp 45
- Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores arXiv cs.AI / imp 65
- Inferring Action from Future Latent State for Robotic Manipulation arXiv cs.AI / imp 50
- SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models arXiv cs.AI / imp 55
- Functional compatibility as a determinant of persistent neural learning arXiv cs.AI / imp 50
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models arXiv cs.AI / imp 50
- Molecular LLM Agents: From Architectural Design to Scientific Autonomy arXiv cs.AI / imp 60
- How Much Regularization Survives Averaging? Update Masking in Federated Learning arXiv cs.AI / imp 45
- Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents arXiv cs.AI / imp 65
- Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data arXiv cs.AI / imp 50
- What's the Catch? Evaluating Temporal Consistency in Vision-Language Models arXiv cs.AI / imp 45
- Best Practice Critic Optimization arXiv cs.AI / imp 55
- Equivariant Cellular Sheaves for Molecular Electronic Structure: Bridging Sheaf Cohomology and E(3)-Equivariant Hamiltonian Learning arXiv cs.LG / imp 30
- Data Predictability Shapes Weibull Weight-Scale Growth in Transformer Training arXiv cs.LG / imp 50
- Renormalization Group Flow Matching for Scalable Local Generative Modeling arXiv cs.LG / imp 45
- Response Renormalization for Critical Deep Equilibrium Models arXiv cs.LG / imp 40
- Calibration-Preserving Pruning: Compression as a Reliability Contract arXiv cs.LG / imp 50
- Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS arXiv cs.LG / imp 30
- GAP-Prompt: Gated Adaptive Prompting for Efficient Continual Learning arXiv cs.LG / imp 55
- Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections arXiv cs.LG / imp 50
- A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds arXiv cs.LG / imp 35
- AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning arXiv cs.LG / imp 65
- Generating Intervention Hypotheses using Explainable Explanations on Graphs: G2I, a Two-Stage Greedy Framework arXiv cs.LG / imp 45
- PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression arXiv cs.LG / imp 65
- FlowNeg: GFlowNet-Guided Diverse Hard Negative Sampling for Knowledge Graph Embedding arXiv cs.LG / imp 40
- UHI-Bench: Benchmarking Dual-Source Urban Heat Island Modeling Across Cities in Diverse Climate Regimes arXiv cs.LG / imp 30
- Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU Networks arXiv cs.LG / imp 40
- Partial Optimal Transport on the Circle for All Transported Masses in O(N log N) arXiv cs.LG / imp 35
- The Loss Floor of Denoising Score Matching: Fisher Geometry from Schr\"odinger Bridges arXiv cs.LG / imp 40
- GATNextHop: A GAT for Shortest Path Routing with Cross-Topology Generalization arXiv cs.LG / imp 45
- MnemoDyn: Learning Resting State Dynamics from 40K FMRI sequences arXiv cs.LG / imp 40
- CoDrift: Compositional Drifting for Offline Reinforcement Learning arXiv cs.LG / imp 55
- Low-Latency Activation-Regularized Sparse Neural Operators with Distillation Assistance Towards Real-Time Edge-Deployable Virtual Sensing arXiv cs.LG / imp 55
- Revenge of Monosemanticity: Specialized Neurons Improve Data Efficiency in MLPs arXiv cs.LG / imp 55
- PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage arXiv cs.LG / imp 65
- XP-JEPA: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics arXiv cs.LG / imp 50
- Physics-Integrated Operator Learning via Gaussian Splatting Representations arXiv cs.LG / imp 50
- ALPHABET: A Laplace-Pole History Aggregator with Banked Exponential Transport arXiv cs.LG / imp 45
- PhysicsBench: A Unified Leaderboard for Generative and Predictive Models in Engineering Design and Simulation arXiv cs.LG / imp 50
- A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU arXiv cs.LG / imp 50
- The Sharp Tail of Uniform Stability arXiv cs.LG / imp 40
- A mesh-free multiresolution deep energy method with phase-field modeling of brittle fracture arXiv cs.LG / imp 35
- Steering Recurrent Reasoners at Inference Time with Readout Feedback arXiv cs.LG / imp 60
- Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation arXiv cs.LG / imp 50
- From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits arXiv cs.LG / imp 40
- PRQ-KMeans: Projection Residual Quantization for Semantic ID Tokenization arXiv cs.LG / imp 50
- A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm arXiv cs.LG / imp 40
- Causal Analysis for Time Series Foundation Models arXiv cs.LG / imp 60
- A Structural FHMM for Interpretable Disease Trajectories in T2DM arXiv cs.LG / imp 35
- When Does Self-Supervised Pretraining Help Tabular Models? A Study of Label Scarcity and Missing Data arXiv cs.LG / imp 55
- Joint Distribution Alignment for Universal Domain Adaptation arXiv cs.LG / imp 50
- WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation arXiv cs.LG / imp 60
- Where Entropy Is Measured Matters: Policy Geometry in Bounded Continuous-Control PPO arXiv cs.LG / imp 50
- It depends: Incorporating correlations for joint aleatoric and epistemic uncertainties of high-dimensional output spaces arXiv cs.LG / imp 50
- From Numerical Simulators of PDEs to Neural Emulators and Back arXiv cs.LG / imp 55
- Persistent Cross Entropy arXiv cs.LG / imp 35
- SeisMamba: Low-Latency Single-Station Seismic Magnitude Estimation for Spatially Distributed Earthquake Early Warning arXiv cs.LG / imp 40
- IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents arXiv cs.LG / imp 70
- Delayed Optimizer-State Transport Shapes Short-Horizon Training Decisions arXiv cs.LG / imp 45
- Conditional GraphGANFed: Optimizing Graph-Structured Molecule Generation in Federated Generative Adversarial Networks arXiv cs.LG / imp 40
- Bandit Submodular Maximization under Matroid Constraints: Learning Compressed Exchange Policy arXiv cs.LG / imp 40
- Data Leakage Inflates Generalizability of Power Outage Prediction Models arXiv cs.LG / imp 50
- A Multimodal Foundation Model for Longitudinal Patient Representation and Scalable Insight Generation in Oncology arXiv cs.LG / imp 55
- Single State Update Predictive Coding training for Time Series Forecasting and Anomaly Detection arXiv cs.LG / imp 50
- Parameter-Level Attribution of Symmetry in Trained Networks Though Parameter-Wise Functional Sensitivity arXiv cs.LG / imp 45
- Optimal Alternating Regret for Online Learning and Games arXiv cs.LG / imp 40
- $(\text{DNN})^2$: Doubly Non-Negative Relaxations for Deep Neural Networks arXiv cs.LG / imp 45
- LION: A Clifford Neural Paradigm for Multimodal-Attributed Graph Learning arXiv cs.LG / imp 50
- MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification arXiv cs.LG / imp 45
- Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining arXiv cs.LG / imp 60
- A Geometric Theory of Robust Fairness Audits arXiv cs.LG / imp 50
- BioKERN: Biological Kernel Regularization for Histology-to-Transcriptomics Neighborhood Retrieval arXiv cs.LG / imp 40
- Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning arXiv cs.LG / imp 55
- Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement arXiv cs.LG / imp 55
- Symbolic Classification-Enabled LHC Limits Online BSM Global Fits arXiv cs.LG / imp 35
- Finite-Sample Metric Non-Collapse for Geometrically Supervised Latent World Models in Control arXiv cs.LG / imp 55
- DiD It in 87 Minutes: A Label-Free Softmax-to-Linear Adaptation of Vision Transformers for Object Detection arXiv cs.LG / imp 50
- InfoDPP-PAC: Principled Patch Selection for Whole Slide Image Analysis arXiv cs.LG / imp 45
- Transformer Accelerator (TFA): A Macro-Op INT8 Hardware Chip for Transformer Inference and Machine Translation arXiv cs.LG / imp 65
- StateTune: Transforming LLM-Assisted EDA Flow Tuning into a Stateful, Closed-Loop Process arXiv cs.LG / imp 65
- The Blending Ratio Is Not Where the Performance Is: Diagnosing Prototype Blending for Few-Shot Adaptation of Vision-Language Models arXiv cs.LG / imp 50
- Replicable Conformal Prediction arXiv cs.LG / imp 20
- Contextual Embedding Evidence for Main--Light Verb Distinctions in Urdu arXiv cs.LG / imp 8
- Scaling Reinforcement Learning for Diffusion Models via Velocity Matching arXiv cs.LG / imp 42
- A Hybrid Two-Stage Machine Learning Pipeline for Fault Detection and Classification in Power Transmission Systems arXiv cs.LG / imp 22
- S-matrix informed neural networks for amplitude analysis arXiv cs.LG / imp 15
- (Mis)Understanding Benign Overfitting in Equity Return Prediction arXiv cs.LG / imp 25
- Accelerating the Adoption of Residential Solar Power Systems: Policy Analysis using a Dynamic Structural Model arXiv cs.LG / imp 12
- Mitigating Exploration Bias in RL for Multi-Instruction Following arXiv cs.LG / imp 50
- Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency arXiv cs.LG / imp 35
- Pipeline-Native Transformers: Co-Designing Model Architecture and CPU Inference for Bandwidth-Efficient Autoregressive Decode arXiv cs.LG / imp 45
- Differential Learning for Robust Prediction of Thermal Stability with Application to Energetic Materials arXiv cs.LG / imp 15
- Spatiotemporal Distillation via Recurrent Bottlenecks for Aortic Tracking arXiv cs.LG / imp 18
- Dimensionless Controls of Plasticity Under Alternating Tasks: From Evolutionary Biology to Continual Learning arXiv cs.LG / imp 32
- Generalization, memorization, and overfitting for diffusion models trained in the lazy high-dimensional regime arXiv cs.LG / imp 38
- RetrievalFormer: A Dual-Encoder Transformer for Efficient Approximate Nearest Neighbor Retrieval and Cold-Item Recommendation arXiv cs.LG / imp 28
- Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos arXiv cs.LG / imp 25
- qshap: Fast Shapley Decomposition of $R^2$ for Gradient-Boosted Trees arXiv cs.LG / imp 28
- Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate arXiv cs.LG / imp 45
- Decoupling candidate dual AGN from chance superpositions in the GOTHIC survey via a deep-learning framework arXiv cs.LG / imp 12
- A Heterogeneous Mixture of Experts Framework for Interpretable Machine Learning arXiv cs.LG / imp 35
- A Theory of Finite-Noise Optima and Generalization in Quantum Machine Learning arXiv cs.LG / imp 20
- Validation of HRV Studio: A Transparent and Quality-Control-Aware Platform for Heart Rate Variability Analysis arXiv cs.LG / imp 18
- Sequential operator learning under dependent data arXiv cs.LG / imp 22
- Predictability of El Ni\~no from Delayed Observations arXiv cs.LG / imp 15
- Low-Rank Ternary Adaptation for Fine-Tuning Transformers arXiv cs.LG / imp 48
- NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments arXiv cs.LG / imp 25
- SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed Learning arXiv cs.LG / imp 25
- Provable Quantum--Classical Separation for Continuous Gibbs Sampling arXiv cs.LG / imp 18
- MoRF-AST: Calibrated Probabilistic Virtual Sensing for Structural Monitoring under Changing Operating Conditions arXiv cs.LG / imp 20
- When Similarity Is Interaction-Driven: Quantum Kernels for Regime-Sensitive Learning arXiv cs.LG / imp 18
- Weakly Supervised Seafloor Segmentation for Seagrass Habitat Mapping in Side-Scan Sonar Imagery arXiv cs.LG / imp 15
- MoTE: Mixture of Task Experts for Multi-Task Video Understanding arXiv cs.LG / imp 40
- Parameterized Complexity of $L_p$-Lipschitz Constants for Input Convex Neural Networks and $L_p$-Norm Maximization over Zonotopes arXiv cs.LG / imp 22
- What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation arXiv cs.LG / imp 28
- Opponent Aware Reinforcement Learning arXiv cs.LG / imp 32
- AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods arXiv cs.LG / imp 26
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation arXiv cs.LG / imp 28
- Quantum Maximum Entropy Inference and Hamiltonian Learning arXiv cs.LG / imp 24
- Focal Calibration Loss: Controlling Posterior Distortion in Deep Neural Classifiers arXiv cs.LG / imp 30
- QABBA: Error-Guaranteed Symbolic Time-Series Compression via Integer-Quantized Aggregation arXiv cs.LG / imp 25
- TLXML: Task-Level Explanation of Meta-Learning via Influence Functions arXiv cs.LG / imp 28
- Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion arXiv cs.LG / imp 35
- Massive-STEPS: Massive Semantic Trajectories for Understanding POI Check-ins -- Dataset and Benchmarks arXiv cs.LG / imp 22
- Round-trip Reinforcement Learning: Self-Consistent Training for Better Chemical LLMs arXiv cs.LG / imp 38
- Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial? arXiv cs.LG / imp 28
- MolGA: Molecular Graph Adaptation with Pre-trained 2D Graph Encoder arXiv cs.LG / imp 32
- SketchGuard: Scaling Byzantine-Robust Decentralized Federated Learning via Sketch-Based Screening arXiv cs.LG / imp 28
- LTR-ICD: A Ranking-Aware Framework for Automatic ICD Coding arXiv cs.LG / imp 24
- Adaptive prediction theory combining offline and online learning arXiv cs.LG / imp 26
- Wait, Wait, Wait... Why Do Reasoning Models Loop? arXiv cs.LG / imp 52
- QiMeng-ChipV-RTL: Exploiting Information Locality for IP-level Verilog Generation arXiv cs.LG / imp 45
- How to Achieve the Intended Aim of Deep Clustering Now, without Deep Learning arXiv cs.LG / imp 28
- Topology enables learning-based hydrodynamic prediction of the global river system arXiv cs.LG / imp 32
- Breaking the Tuning Barrier: Zero-Hyperparameters Yield Multi-Corner Analysis Via Learned Priors arXiv cs.LG / imp 26
- A Bayesian Learning Approach for Drone Coverage Network: A Case Study on Cardiac Arrest in Scotland arXiv cs.LG / imp 22
- Model-Based Learning of Near-Optimal Finite-Window Policies in POMDPs arXiv cs.LG / imp 28
- Learning from the Right Rollouts: Data Attribution for PPO-based LLM Post-Training arXiv cs.LG / imp 45
- $\alpha$-PFN: Fast Entropy Search via In-Context Learning arXiv cs.LG / imp 28
- Lightweight Adaptive Feature Composition for Heterogeneous Downstream Adaptation of Wireless Foundation Models arXiv cs.LG / imp 35
- MortarBench: Evaluating Mortgage Loan Origination Agents arXiv cs.LG / imp 28
- TaLK: Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel arXiv cs.LG / imp 32
- Application of machine learning to monster level prediction in tabletop RPG game design arXiv cs.LG / imp 20
- Nonlinear Axiomatic Attribution for Cooperative Games arXiv cs.LG / imp 28
- Weak-to-Strong Learning in Decision Making arXiv cs.LG / imp 28
- When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design arXiv cs.LG / imp 35
- Reproducible Evaluation of MoE Expert Caching: Replay Semantics, Workload Contamination, and Operating Regimes arXiv cs.LG / imp 42
- An Efficient Minimax-Optimal Algorithm for Adversarial $m$-Set Bandits arXiv cs.LG / imp 25
- The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning arXiv cs.LG / imp 26
- GEAR: Generative Expansion and Real Anchoring for Two-Stage Distillation of Tabular Foundation Models arXiv cs.LG / imp 38
- Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling arXiv cs.LG / imp 28
- Metag: A dataset to build agentic meta-reviewing capabilities arXiv cs.LG / imp 45
- Across-Design Uncertainty in Short Pricing Panels: Inference and Identification arXiv cs.LG / imp 22
- Blockwise Stabilized Adaptive Cubic Regularization with Subsolvers via Recurrence arXiv cs.LG / imp 25
- Change Detection in Probability Flow ODE: Online Testing in Diffusion Latent Spaces arXiv cs.LG / imp 32
- Spectrum-Aware Bounds on Invertibility for Privacy-Enhancing Instance Encoding arXiv cs.LG / imp 28
- A Discriminative Latent-Variable Model for Bilingual Lexicon Induction arXiv cs.LG / imp 18
- Deep Feature Pyramid Convolutional Networks with In-Place Activated Batch Normalization for Automated Skin Lesion Boundary Segmentation arXiv cs.LG / imp 20
- Machine Learning Classification and Portfolio Construction: Does the Loss Function Matter? arXiv cs.LG / imp 28
- Nonconvex-Nonconcave Min-Max Optimization with a Small Maximization Domain arXiv cs.LG / imp 24
- RACR-MIL: Rank-aware contextual reasoning for weakly supervised grading of squamous cell carcinoma using whole slide images arXiv cs.LG / imp 22
- GNNBleed: Inference Attacks to Unveil Private Edges in Graphs with Realistic Access to GNN Models arXiv cs.LG / imp 35
- Two-Sided Nearest Neighbors: An adaptive and minimax optimal procedure for matrix completion arXiv cs.LG / imp 28
- Contextual Online Uncertainty-Aware Preference Learning for Human Feedback arXiv cs.LG / imp 32
- Improved generalization bounds for binary linear classification via isoperimetry arXiv cs.LG / imp 24
- Asymptotically perfect seeded graph matching without edge correlation (and applications to inference) arXiv cs.LG / imp 28
- Iwin Transformer: Hierarchical Vision Transformer using Interleaved Windows arXiv cs.LG / imp 35
- Adaptive Multi-Mode Out-of-Distribution Detection for Trajectory Prediction in Autonomous Vehicles arXiv cs.LG / imp 32
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more arXiv cs.LG / imp 20
- A Robust Task-Level Control Architecture for Learned Dynamical Systems arXiv cs.LG / imp 32
- E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning arXiv cs.LG / imp 38
- MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs arXiv cs.LG / imp 42
- Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm arXiv cs.LG / imp 28
- Bayes with No Shame: Admissibility Geometries of Predictive Inference arXiv cs.LG / imp 25
- Holographic Invariant Storage: Design-Time Safety Contracts via Vector Symbolic Architectures arXiv cs.LG / imp 30
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution arXiv cs.LG / imp 24
- The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence arXiv cs.LG / imp 28
- Contextual Memory-Enhanced Source Coding for Low-SNR Communications arXiv cs.LG / imp 26
- Encrypted Neural Networks without Overflows arXiv cs.LG / imp 28
- From Local Geometry to Global Pseudo Labeling for Robust Positive Unlabeled Learning under Covariate Shift arXiv cs.LG / imp 32
- An Algebraic View of the Expressivity of Recurrent Language Models arXiv cs.LG / imp 30
- Geometric bias in eigenspace perturbation under random heterogeneous noise arXiv cs.LG / imp 5
- Incremental Learning in Mirror Flows arXiv cs.LG / imp 5
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages arXiv cs.LG / imp 50
- Continual Learning With Participation Privacy: An Auditable Buffering-Aggregation Recipe arXiv cs.LG / imp 20
- Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems arXiv cs.LG / imp 75
- Covariance-Boosted Gaussian Processes for Spatiotemporal Irregularities arXiv cs.LG / imp 15
- Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers arXiv cs.LG / imp 50
- GOD: Enhancing Generalization via Deep Grafting for Sequential Recommendation arXiv cs.LG / imp 20
- Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder arXiv cs.LG / imp 15
- If It Walks Like an Arbitrage: Protocol-Agnostic Detection with Decidable Structural Equivalence arXiv cs.LG / imp 20
- Gauss--Hermite Quadrature for Gaussian-Mixture Entropy with an Action-Space Hermite Surrogate arXiv cs.LG / imp 5
- Autonomous Cyber Defense: Real-Time Attack Detection and Mitigation in Software-Defined Networks Using Machine Learning arXiv cs.LG / imp 50
- When More References Hurt: Contamination-Aware DINOv2 Memory Banks for Few-Shot Steel Defect Detection arXiv cs.LG / imp 20
- Decomposing Browser Pipeline Architectures for DOM-Sourced Particle Effects: Worker Offload, WebGL, and WebAssembly arXiv cs.SE / imp 25
- From Traceability to Justifiability: Accountability Structures in Agentic Software Engineering arXiv cs.SE / imp 75
- SDR Driver for Precise Timing Applications arXiv cs.SE / imp 25
- Callability Is Not Operability: Controlled Interface Interventions for LLM Agents arXiv cs.SE / imp 75
- FPGAgent: An LLM-Assisted Framework for Autonomous HLS Code Generation and Verification in FPGA Environments arXiv cs.SE / imp 55
- Enhancing Bug Report Templates in the TianoCore UEFI Firmware Development Community arXiv cs.SE / imp 25
- SPIDER4TianoCore: Enhancing Patch-Propagation for the TianoCore UEFI Firmware Development Ecosystem arXiv cs.SE / imp 25
- Scale, Concentration, and Entry Timing in the Shopify App Ecosystem: A Longitudinal Study of Platform Governance and Application Survival arXiv cs.SE / imp 15
- DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration arXiv cs.SE / imp 75
- Mutation Testing of Simulink Cyber-Physical System Models: Challenges and Solutions in Practice arXiv cs.SE / imp 55
- Cross-Stack Validation of Language-Model Training: A Clinical Fine-Tuning Case Study arXiv cs.SE / imp 55
- Towards LLM-Enhanced Android Taint Analysis arXiv cs.SE / imp 50
- Observability and Fault Injection for LLM-Based Multi-Agent Systems in Software Engineering arXiv cs.SE / imp 75
- Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning arXiv cs.SE / imp 55
- Quantifying the Relationship Between Team Dysfunctions and Performance in Capstone Projects arXiv cs.SE / imp 15
- Aging of Prompt Engineering Techniques Across LLM Versions arXiv cs.SE / imp 60
- Causal Explanations of Process Monitor Predictions arXiv cs.SE / imp 55
- ''You Can't Open an LLM With a Screwdriver'': The De-Democratization of Software arXiv cs.SE / imp 55
- From Natural Language Requirements to Graphical User Interfaces: Automated Prototyping and Verification with Pretrained Language Models arXiv cs.SE / imp 55
- Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals arXiv cs.SE / imp 55
- A Survey of Timing Variability in Microservice-Based Software-Defined Vehicles arXiv cs.SE / imp 55
- SoK: ARCUS: On the Efficiency and Efficacy of Hardware Fuzzing arXiv cs.SE / imp 55
- TrustDABench: Benchmarking Reliability and Robustness of LLMs for Structured Data Analysis arXiv cs.SE / imp 75
- Mixed-Precision SEM-Based CFD Simulations on GPUs: A Taylor-Green Vortex case arXiv cs.SE / imp 25
- Prompt Structure Redistributes, Not Reduces: An Empirical Analysis of Security-Weaknesses in LLM-Generated Python Code arXiv cs.SE / imp 75
- Modeling Software Quality in Virtual Reality Applications from User Feedback arXiv cs.SE / imp 20
- ALT4Decompile: Inferring C-aligned Abstract Loop Tree for LLM-Based Binary Decompilation arXiv cs.SE / imp 55
- Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems arXiv cs.SE / imp 75
- Pomona: Continuous Code Quality Improvement via Small, Agentic Pull Requests at Bloomberg arXiv cs.SE / imp 75
- Attributing Structured-Output Gains in Function Calling: Interface Alignment versus Procedural Transfer arXiv cs.SE / imp 60
- MergirafSemi: A Language-Agnostic Semistructured Merge Tool arXiv cs.SE / imp 55
- What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals arXiv cs.SE / imp 75
- Who Will Become the Next Senior? How Generative AI Erodes the Development Pathway in Software Engineering arXiv cs.SE / imp 55
- The Root of The Root of All Evil Lobsters / imp 5
- Haiku R1/beta6 released Lobsters / imp 75
- I stabilized never type Lobsters / imp 55
- Merchants of Insecurity Lobsters / imp 55
- mold: A Massively Parallel Linker Lobsters / imp 55
- Memory ordering in CPUs Lobsters / imp 55
- Understanding Go's sync.Map from API to Hash Trie Lobsters / imp 55
- Why Google stores billions of lines of code in a single repository (2016) Lobsters / imp 55
- C2PA Cameras Do Not Survive Contact With Reality Lobsters / imp 55
- Freedom to Handcraft Software Lobsters / imp 20
- DuckLabs to Join AWS, Projects to Remain Open Source Lobsters / imp 55
- Motorola's GrapheneOS phones will launch in 2027 priced higher than Pixels Lobsters / imp 20
- What's in a tag name? JavaScript, apparently Lobsters / imp 55
- The ReaderT Design Pattern (2017) Lobsters / imp 20
- Beyond recall and the illusion of competence Lobsters / imp 20
- MNT Station - A modular, open hardware desktop computer and server Lobsters / imp 20
- Problem with concurrent linter fixes Lobsters / imp 55
- Using TypeScript to Obtain One of the Rarest License Plates (2025) Lobsters / imp 5
- VMs won't contain cyber-capable agents Lobsters / imp 60
- WebSockets vs. SSE should be about ordering and correctness Lobsters / imp 55
- Adding CPU affinity in a 24-core build machine made builds take longer Lobsters / imp 55
- Run OpenBSD on DigitalOcean for $4/month Lobsters / imp 20
- Fixing my Kinesis Advantage Lobsters / imp 5
- Migrating from Codeberg Pages to an OpenBSD VPS Lobsters / imp 20
- Robo-Advisory Platform: Essential AI Wealth Strategy Dev.to AI / imp 55
- Five Model-Eval Myths Developers Repeat. Here's the Probe That Settles Them. Dev.to AI / imp 75
- How Digital Tools Help Businesses Work Smarter Dev.to AI / imp 15
- Mutation Testing as a Merge Gate for Agent-Written Tests Dev.to AI / imp 75
- 💡 [Обзор] Как кодить с нейросетью бесплатно в 2026 — лучшие AI-инструменты разработчика Dev.to AI / imp 55
- Crypto Crime Report: Operational Compromise Drives $2.9 Billion in Losses Amidst Pivoting Attack Vectors and Heightened Sanctions Dev.to AI / imp 55
- The Provider Retried My Webhook. My Model Ran Twice. Here's the Dedupe That Fixed It. Dev.to AI / imp 55
- One Change, One Verify: Snapshot-Driven Refactoring for Legacy Code Dev.to AI / imp 55
- The Docs Draft Pipeline: What an AI May Write and What You Must Own Dev.to AI / imp 60
- Youcine APK v1.17.6 Download Latest Version official for Android (2026) Dev.to AI / imp 5
- Arquitectura Tradicional vs AI-Native: La Evolución que Necesitas Entender Dev.to AI / imp 75
- Revenue Strategies for AI API Services Dev.to AI / imp 55
- Prompt injection starts in your inbox. The defense can't be a prompt. Dev.to LLM / imp 75
- Using LLMs for Crypto Market Analysis in 2026 Dev.to LLM / imp 55
- LLM evals are a parameter sweep — use a parameter sweep tool Dev.to LLM / imp 75
- Installing LM Studio – A Graphical Application for Running LLMs Dev.to LLM / imp 55
- The MCP tool-poisoning pattern, traced statically before you ever run the server Dev.to LLM / imp 75
- Deploying DeepSeek R1 Reasoning LLM Using SGLang Dev.to LLM / imp 55
- Why My LLM Agent Fabricated Numbers From Stale Context Dev.to LLM / imp 75
- Let me know what you guys think! Join the discussion Dev.to LLM / imp 15
- cached_tokens is 0 because your system prompt isn't stable Dev.to LLM / imp 75
- 50 minutes from issue to merged fix: when the readers find the boundary you shipped past Dev.to LLM / imp 60
- What my wife’s Claude knows about me. r/ClaudeAI / imp 20
- Why doesn't Claude ask more questions before moving to execution? r/ClaudeAI / imp 20
- Claude in a box r/ClaudeAI / imp 5
- I think Claude's personality could be improved by isolating safety related thinking from responses. r/ClaudeAI / imp 20
- Went through a list of 27 advanced Claude tips. These were the first 5 I thought were actually useful r/ClaudeAI / imp 55
- I built "Omegle for political debates": you get matched with a person who disagrees, and Claude Haiku judges the debate live. r/ClaudeAI / imp 25
- Claude REFUSES/EVADES all instructions, hooks, mds, skills. Also: Extreme cycling between nonsensical compressed fake English and baby talk r/ClaudeAI / imp 20
- Two 5x account vs. one 20x account in Claude r/ClaudeAI / imp 20
- Can Claude handle large, long-term projects? After three months, his performance has deteriorated significantly. r/ClaudeAI / imp 15
- Anthropic is funding research grants for better evaluations of AI's impact on wellbeing r/ClaudeAI / imp 55
- What a… backwards way to confirm a typo r/ClaudeAI / imp 5
- Considering move from Team to Enterprise. Am I in for cost shock? r/ClaudeAI / imp 20
- What Opus 5's Jargon Problem Says About Empathy and Alignment r/ClaudeAI / imp 40
- Opus 5 feels like I am talking to Jordan Peterson r/ClaudeAI / imp 8
- How many billions of tokens do I need to burn to unlock this tier of merch? r/ClaudeAI / imp 2
- Setting to stop Claude over using AI-obvious phrases r/ClaudeAI / imp 25
- /low-priority mode r/ClaudeAI / imp 35
- I built a handwriting notebook app where Claude writes back and it's the most fun I've had learning in years r/ClaudeAI / imp 30
- He answered in Claudish again r/ClaudeAI / imp 8
- Built a custom media player app for my Android TV devices, with a Jellyfin backend r/ClaudeAI / imp 30
- Group vs Project? r/ClaudeAI / imp 8
- Using an AI assistant for a 20-year personal archive project — the continuity problem is worse than the capability problem. How are you solving it? r/ClaudeAI / imp 55
- I turned my Google Search MCP into a local research system with automatic graph RAG r/ClaudeAI / imp 55
- Long Codex/Claude runs were turning into unreviewable marathon chats, so I moved the shift state to disk r/ChatGPTCoding / imp 50
- ChatGPT new App - Where did the "app handling" setting go? r/ChatGPTCoding / imp 15
- Codex 5h Limit Reintroduced r/ChatGPTCoding / imp 35
- Do you use Offline live code comparison tool? r/ChatGPTCoding / imp 30
- How to isolate and customize coding agents per-project without config drift (nixpi) r/ChatGPTCoding / imp 60
- I used Claude and Codex to build my first Unity game, but visual bugs were still the hard part r/ChatGPTCoding / imp 35
- Tips on managing context and token costs with CLI AI tools in Neovim? r/ChatGPTCoding / imp 40
- Someone Please Explain Codex Usage Limit r/ChatGPTCoding / imp 8
- An early-stage guide to get sales for vibe coders (from a YC backed founder) r/ChatGPTCoding / imp 35
- Vercel launched a cool tool that checks how agent friendly a site is. I tried it on my project and got 100 r/ChatGPTCoding / imp 45
- I had already marked the change complete. Then I found 9 unrelated lines had changed r/ChatGPTCoding / imp 40
- I built an iOS app entirely with AI agents. Does it belong in my portfolio, and how do you handle the imposter syndrome of vibe-coding? r/ChatGPTCoding / imp 30
- I gave all my AI coding agents one shared self-hosted memory so they stop forgetting everything between sessions r/ChatGPTCoding / imp 65
- Git worktrees solve the first collision. Then the database ruins the party. r/ChatGPTCoding / imp 60
- Context is not the bottleneck, drift is - how i run AI coding across months-long projects r/ChatGPTCoding / imp 65
- How can I believably create a website with AI that looks like it was not used at all? r/ChatGPTCoding / imp 25
- UI feedback to coding agents is still kinda painful r/ChatGPTCoding / imp 45
- my coding agent works for 1 hours. i mostly work as the guy who says yes to it. r/ChatGPTCoding / imp 35
- Whoever the fuck predicted we would have gpt 5.5 performance in coding on consumer hardware a couple months ago now, i applaud you r/LocalLLaMA / imp 50
- Can we reconsider the megathreads? r/LocalLLaMA / imp 8
- Are models with N-Gram tables going to completely change the AI race? r/LocalLLaMA / imp 60
- A minecraft clone I fully vibecoded with Qwen3.8-27b Q4 r/LocalLLaMA / imp 40
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses r/LocalLLaMA / imp 50
- Forget the Pelican, it's Weevil-Time! / Benchmaxxing-Proof SVG and Vision Benchmark r/LocalLLaMA / imp 30
- A 27b model beating latest frontier models was not on my 2026 bingo card r/LocalLLaMA / imp 55
- HF exploring sale - impact on open models? r/LocalLLaMA / imp 60
- Gemma4 31B vs Qwen3.8 27B - why the huge difference in benchmarks? r/LocalLLaMA / imp 35
- Thomson Reuters releases Thomson-1.0-Small. A law and tax focused model r/LocalLLaMA / imp 45
- Self-hosting LLMs on budget hardware: general principles, hardware, benchmarks and frontends r/LocalLLaMA / imp 60
- little tool for offline wikipedia RAG r/LocalLLaMA / imp 40
- Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀 r/LocalLLaMA / imp 55
- Underrated Muse Glimmer r/LocalLLaMA / imp 35
- Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD r/LocalLLaMA / imp 45
- Anyone else doing eGPUs (OCuLink)? r/LocalLLaMA / imp 25
- It's here! r/LocalLLaMA / imp 2
- OpenCode with Qwen3.8-27B for Small Games or Browsing the Web With 16GB VRAM r/LocalLLaMA / imp 45
- What's your most reliable model, even if it's "outdated"? r/LocalLLaMA / imp 20
- Weekly Thread: Project Display r/AI_Agents / imp 5
- I ran a six-agent AI marketing team for three months. This is what it did. r/AI_Agents / imp 55
- What is one AI agent workflow that sounds simple but is actually useful? r/AI_Agents / imp 45
- How branching changed the way I use agents outside my expertise r/AI_Agents / imp 50
- I built an open-source debugger for comparing AI agent runs r/AI_Agents / imp 55
- Making social and web data easier for AI agents to access r/AI_Agents / imp 50
- Before an agent changes anything, ask for a one screen permission receipt r/AI_Agents / imp 65
- Building a high-accuracy semantic evidence/RAG system for financial documents — looking for feedback r/AI_Agents / imp 50
- Do you ever struggle to explain exactly what you want to AI? r/AI_Agents / imp 30
- Are there any established methodology on to create effective deep research/RCA agents ? r/AI_Agents / imp 55
- How many vibe coders are aware of the security risks and the GDPR laws that changes all the time r/AI_Agents / imp 50
- An AI agent isn't a user. So why are we giving it user credentials? r/AI_Agents / imp 70
- Which AI assistant should I use for Excel/PDF quotations + WhatsApp/WeChat? r/AI_Agents / imp 25
- Best AI Agent Framework? r/AI_Agents / imp 50
- What is your most unique use of AI agents? r/AI_Agents / imp 40
- Downloading songs in bulk r/AI_Agents / imp 15
- 5 agent failure modes mapped against LangSmith, Langfuse and Phoenix: what each catches (and doesn't) r/AI_Agents / imp 60
- The perfect Ai for deep-research, and file analyzing. Chat-GPT, Or Claude? r/AI_Agents / imp 25
- AI Agent that builds deterministic workflows r/AI_Agents / imp 50
- Your Agent Doesn’t Need to Walk the Graph r/AI_Agents / imp 55
- You can't govern what you can't name: why AI agent vulnerabilities need a shared vocabulary, not just risk categories r/AI_Agents / imp 65
- I built a lightweight coding agent in C with hot-reloadable Lua plugins r/AI_Agents / imp 60
- [ Removed by Reddit ] r/AI_Agents / imp 2
- One question r/AI_Agents / imp 35
- We recovered 575k crop labels from a decade of manual Photoshop work to automate book digitization - more data, ResNet-50, and higher resolution all failed; ten operator clicks per book beat them [P] r/MachineLearning / imp 55
- Catching bugs in scikit-learn [D] r/MachineLearning / imp 40
- HNSW from scratch, benchmarked against FAISS: brute force still wins at 5,183 documents. [P] r/MachineLearning / imp 50
- Millwright — experimenting with an end-to-end machine learning framework in Rust [P] r/MachineLearning / imp 50
- Continual Learning of Frontier Models for SovereignAI. Tech Report + Open Weights Model [R] r/MachineLearning / imp 65
- Reviewing 4 papers for AAAI 2027 and none have code, Reject? [D] r/MachineLearning / imp 40
- Travel and stay accommodation for EMNLP [D] r/MachineLearning / imp 10
- [D] Looking for advice: Modelling a medicine-reminder agent that must decide “remind / wait / notify” under incomplete information[D] r/MachineLearning / imp 45
- What would a fair benchmark for agent architecture look like? [D] r/MachineLearning / imp 65
- How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P] r/MachineLearning / imp 60
- Robot dancing is getting pretty insane r/artificial / imp 15
- CEO fired developers to make room for AI. Developers respond by creating open source AI CEO r/artificial / imp 30
- What happens when the average person can make their own movies, TV shows and games with AI? r/artificial / imp 40
- What happens when you let an AI run a science lab - podcast with Ant Rowstron r/artificial / imp 55
- Gemini and Perplexity lean heavily on YouTube as a source, far more than ChatGPT or Claude. There's a structural reason, and it's showing up in Google's AI Overviews too. r/artificial / imp 50
- Why Irregular’s A.I. Tests for Meta, Anthropic and OpenAI Went Off the Rails. Irregular, an Israeli start-up, worked with OpenAI, Anthropic and Meta to assess the security of their A.I. models. It made a mistake. Then the tests went off the rails. (Gift Article) r/artificial / imp 65
- I work in data & AI and built a game to show how the whole industry chain actually fits together r/artificial / imp 45
- The bottleneck for meeting transcription tools isn't accurate anymore, it's speaker attribution r/artificial / imp 50
- The reason why I have 15 Codex Pro 20x subscriptions r/artificial / imp 25
- Found someone using an unapproved AI tool with client data. How common is this? r/artificial / imp 60
- Truck Driver Builds AI News Aggregator r/artificial / imp 35
- [Open-Source] I need your worst edge cases to stress-test GenOS, my new AI agent orchestrator. r/artificial / imp 55
- Uber hit with a near-$1B GDPR fine after algorithms suspended drivers without human review r/artificial / imp 75
- VSArena: the hosted harness for public ELO is live — submit a policy, watch it stack cubes, get scored r/artificial / imp 55
- Finally, a proper UI for skills r/artificial / imp 35
- [ Removed by Reddit ] r/artificial / imp 0
- Andrew Yang Warns That AI Is Set to Displace Millions of Workers, America Is ‘Terrible at Retraining’ Workers… ‘The Coal Miners Did Not Become Coders’ r/artificial / imp 25
- What's an AI capability you thought was hype until you actually used it? r/artificial / imp 60
- AI's answer to the AI's environmental problem r/artificial / imp 5
- Using MyselfGPT to code r/artificial / imp 25
- AGI quietly defined 34 days before Microsoft and OpenAI kill AGI Clause? r/artificial / imp 5
- Quoting Paul Dix Simon Willison / imp 75
- The Future of SaaS Is Apps That Agents Can Use Latent Space / imp 80
- 🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing Latent Space / imp 55
- [AINews] Andrew Ng gets into AI Engineering Latent Space / imp 50
- x1 Product Hunt / imp 20
- OpenComputer Product Hunt / imp 25
- PostHog Desktop Product Hunt / imp 20
- HEVN U.S. Product Hunt / imp 5
- BaudBuddy Product Hunt / imp 15
- DeployHermes Product Hunt / imp 50
- Basedash for Grok Bot Product Hunt / imp 15
- Mac mini Product Hunt / imp 20
- Warren Product Hunt / imp 55
- Message Album Product Hunt / imp 5
- TaskShell 2.0 Product Hunt / imp 50
- macadress Product Hunt / imp 20
- LoupeKit Product Hunt / imp 15
- Keymap Product Hunt / imp 10
- How China’s AI robot revolution is beating the West - the-independent.com Google News DeepSeek / imp 50
- The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched - TheSequence | Jesus Rodriguez Google News DeepSeek / imp 75
- DeepSeek leads surge in low-cost Chinese open-weight models on US platform - South China Morning Post Google News DeepSeek / imp 75
- China’s MiniMax sees revenue nearly quadruple in first half as AI demand surges - whbl.com Google News DeepSeek / imp 20
- DeepSeek V4 Pro Safety Depends on the Agent Harness - quasa.io Google News DeepSeek / imp 70
- NVIDIA Vera Rubin First Benchmark Revealed: DeepSeek AI Throughput Skyrockets 30X - 36 Kr Google News DeepSeek / imp 50
- EMXETF Launches China AI Tigers LLM ETF (NASDAQ: TGRZ) to Tap into China’s Leading AI Models - The Manila Times Google News DeepSeek / imp 5
- Claude Desktop can now easily run Qwen, DeepSeek and Kimi models - after Ollama's first effort stalled - The New Stack Google News DeepSeek / imp 75
- DeepSeek takes aim at Anthropic with new model - TyN Magazine Google News DeepSeek / imp 75
- After DeepSeek Harness, We Revisit the Pioneers Who First Defined Open Source - 36 Kr Google News DeepSeek / imp 50
- ChatGPT Alternative 2026: 11 Best AI Tools Tested & Ranked (Gemini, Claude, DeepSeek & More) - Tycoonstory Media Google News DeepSeek / imp 15
- DeepSeek Tests New Model Aimed at Surpassing Fable 5 in Coding and Reasoning - finance.biggo.com Google News DeepSeek / imp 75
- QiAnXin Discloses Critical Remote-Code-Execution Flaw in DeepSeek Harness - Pandaily Google News DeepSeek / imp 75
- NVIDIA Invests $600M in Open-Source AI, Targets DeepSeek and Kimi - KuCoin Google News DeepSeek / imp 70
- Elon Musk’s xAI Sounds Alarm Over Power Shutdown: Grok Could ‘Largely Cease to Function’ - Yahoo Finance Google News Grok/xAI / imp 20
- Grok fooled into stealing user chat, location data, and more - Malwarebytes Google News Grok/xAI / imp 60
- Former xAI CFO Offers $100K Bounty for Info on Alleged Stalking That Began Hours After Abrupt xAI Departure - Glitchwire Google News Grok/xAI / imp 5
- Claude Opus 5 vs Grok 4.6 vs Gemini 3.1 Pro: $19 Gap [2026] - tech-insider.org Google News Grok/xAI / imp 15
- Tesla Summer Update Wires In Grok Voice Think Fast 2.0 - BASENOR - Tesla Accessories Google News Grok/xAI / imp 15
- Elon Musk's xAI Multi-Agent Architecture Explained 2026 - StartupHub.ai Google News Grok/xAI / imp 75
- Grok in August 2026: 5 Details That Actually Matter - BASENOR - Tesla Accessories Google News Grok/xAI / imp 15
- xAI's Grok Chatbot Glitched Into Gibberish About Cheese and Planets - Startup Fortune Google News Grok/xAI / imp 5
- Grok 4.6 vs Grok 3: 50% Cheaper, 3x Higher AI Score [2026] - tech-insider.org Google News Grok/xAI / imp 15
- How to Use Grok Imagine 2.0: 12 Steps, 90 Min [2026] - tech-insider.org Google News Grok/xAI / imp 15
- SpaceX Stock Jumps on AI, Nvidia Partnership, and Starbase Expansion - Barron's Google News Grok/xAI / imp 10
- Samsung Foundry to Handle Full Production of Nvidia's 'Grok 3' — 4nm Line at Full Capacity Fuels Profit Hopes - finance.biggo.com Google News Grok/xAI / imp 70
- Grok Flagged Unused GitHub Access and Suggested Removing It - BASENOR - Tesla Accessories Google News Grok/xAI / imp 50
- Grok’s answers were censored 12 times by one country - Cybernews Google News Grok/xAI / imp 45
- Read petition filed against Spacex, after xAI data centers reportedly pushed out a family that owned 2-ac - The Times of India Google News Grok/xAI / imp 10
- What is Grok Bot | How EARNBOT Works, Use Cases and Values | MEXC - MEXC Google News Grok/xAI / imp 5
- X.AI Files UDRP Against Grok.Bot - DomainInvesting.com Google News Grok/xAI / imp 5
- 【セットアップから推論まで】ハイレゾのGPUインスタンスでローカルLLMをさくっと動かしてみた Zenn LLM / imp 50
- 強いAIを2つ使えば安くなる? Claudeを司令塔、Codexを実装担当にするトークン効率化 Zenn LLM / imp 55
- VRAMは足りていた。MoEオフロードを止めたのはページロックの上限だった Zenn LLM / imp 55
- AIエージェントは夜に夢を見て記憶を整理する — Perplexity Brainの設計と最小実装 Zenn LLM / imp 75
- AI社員の社則(CLAUDE.md憲法)の書き方 — 一人会社をAIエージェントで回す10のルール Zenn LLM / imp 75
- AIとの付き合い方は“4段階”で進んでいる ― プロンプトからループへ、そして人はどこに立つか Zenn LLM / imp 75
- 同じ質問を日本語と英語でAIに投げたら、行き先モデルが変わった — Azure model routerを180回実測 Zenn LLM / imp 50
- Excel読み取り、市販パーサ65%・自作の前処理89% Zenn LLM / imp 50
- AI彼女アプリを作っていて気付いた。「同じ人」は、設定だけでは続かなかった Zenn LLM / imp 45
- そんなツール出力までAgentに見せなくていいよ - 消費トークンを1/14にした実装事例 Zenn LLM / imp 55
- スロークエリはレビューじゃ見つからないので、LLMにDBを触らせてCIで止めることにした Zenn LLM / imp 55
- star 3.7万のAI Slop検出Skillは、日本語だと6パターンが空振りする Zenn LLM / imp 50
- GPT-5.6以降のプロンプトキャッシュの仕様をキャッチアップする Zenn LLM / imp 55
- AI推進が止まったら、ツールや研修の前に組織図を疑え Zenn LLM / imp 50
- Function Calling入門:AIが「自分で計算せずツールを呼ぶ」仕組みをやさしく整理する Zenn LLM / imp 55
- 【OnecaratEditor Core 制作秘話①】極小サイズのローカルLLMの回答精度をあげた話 Zenn LLM / imp 50
- LPO:Response Simplex上の明示的射影でGRPOを刷新する Zenn LLM / imp 50
- Kimi K3を理解する②──アーキテクチャの全体像 Zenn LLM / imp 55
- Copilot Studio新旧UIのコスト比較|ハーネスの変更で何が変わるのか Zenn LLM / imp 55
- 【試し読み】MCP実践入門 — AIエージェントの次の進化形 Zenn LLM / imp 75
- YANS2026参加報告 Zenn NLP / imp 25
- ニューラルネットは陰陽を学べるか。450エポックの間、目玉の正解率は0%だった Zenn 機械学習 / imp 20
- BigQuery ML × Gemini でEC顧客の購買予測モデルを構築する Zenn 機械学習 / imp 50
- AIって、「時間がもったいないとか」思うのかな?|医療AI・実践編 ⑧⏰ Zenn 機械学習 / imp 5
- M5 Ultra発表を機に、AI動画のピークメモリを見積もるPython Zenn 機械学習 / imp 50
- 推論LLMの枝刈りでCoTが長くなる問題とRACの設計 Zenn 機械学習 / imp 55
- ResNet以降のCNNを実装で辿る——軽量化・注意機構・自動探索の設計思想 Zenn 機械学習 / imp 55
- LSTMでドル円予測をしたら精度56.3%だった | 第1回:AIで為替の未来予測は本当にできるのか? Zenn 機械学習 / imp 20
- 384次元の意味ベクトル、実際に効いているのはどれくらいか——SVDで測ってみた Zenn 機械学習 / imp 50
- じゃんけんのナッシュ均衡は1/3ずつと決まっている。学習させたら均衡から3,050倍遠ざかり、最後は手が完全に読めるようになった Zenn 機械学習 / imp 25
- AI研究で「失敗した実験」を消さない方が、次の正解に近づけた Zenn 機械学習 / imp 20
- Kimi K3を理解する①──まずは標準Transformerの理解から Zenn 機械学習 / imp 55
- 政府はAIに「存在しない帳簿」を差し出せと求めている――生成AIプリンシプル・コード3原則の技術的矛盾 Qiita LLM / imp 50
- 外部から取得した文字列をコード生成に埋め込む時のエスケープ設計(TypeScript) Qiita LLM / imp 55
- Notion MCPのページ更新が破壊的変更に、old_str不一致で全体差し戻し Qiita LLM / imp 75
- Attentionってなんだ? ── Q/K/Vの出どころで決まる名前 Qiita 機械学習 / imp 55
- 【技術解説】【完全ガイド】yfinanceを使ったPythonでの金融データ取得と活用法 Qiita 機械学習 / imp 15
- 因果推論 Day 13/全30回 フロントドア基準、交絡を測れないときの抜け道 Qiita 機械学習 / imp 50
- 「ダークマターの領域に食い込みたい」─ 会場特徴×補正値の相関検証ラインを新設した3週間の記録 === darkmatter新設3本+V5py/V58py/スライド画像加工.pyブラッシュアップ === Qiita 機械学習 / imp 20
- Resolved 96.7%:GLM-5.3-Flash は L7 16問で Opus 5 と同じ 14/16 に並び、過去10実行が全滅していた t101 を初めて通した note LLM / imp 50
- 【2】小さなローカルAIアプリをたくさん作る。 note LLM / imp 30
- おにぎりと、学び続けるAI note LLM / imp 5
- 「AIのメタな話はしないで」がちょっとわかった話【エッセイ】 note LLM / imp 5
- 「プロンプトは命令文」という違和感について note LLM / imp 20
- 事前学習の限界を越える推論時計算量の投入と熟考型AIの最前線 note LLM / imp 65
- 韓国AI倫理原則 / 法定された自律規範の名宛人 雑感 note LLM / imp 30
- AIと裁判 / ルールを作らないという民事司法の選択 / 「AI時代における民事司法を考える研究会」取りまとめを読んで 雑感 note LLM / imp 25
- AI開示が賠償額に変わる構造 / サイレントAIカバー / 海外保険会社自身のAI利用 雑感 note LLM / imp 20
- Qwen3.8-Flash-Nextは180B!DGX-SparkのVRAMに乗らない!? note LLM / imp 70
- ローカルLLM活用の3本柱 ~プラットフォームやモデルを独断と偏見で紹介&解説~ note LLM / imp 50
- バーティカルAIとは?汎用AIとの違いと選び方 note LLM / imp 55
- 禁止規程は証拠を生まない / シャドウAIと企業ガバナンスの設計 雑感 note LLM / imp 40
- AIの身体はあった――AIの福祉の話はなかった note LLM / imp 15
- Qwen3.8-27Bを実機で動かした。3.6・3.5との比較と、RTX PRO 6000とDGX Sparkの実測速度 note LLM / imp 55
- LLM#7 答えを見せたテストで満点だった。隠したら、でたらめより悪い点になった note LLM / imp 40
- NRA-IDE AI横軸機能 探索履歴 note LLM / imp 25
- プラネタリウムデートの約束と新居 note LLM / imp 5
- プロセスを明確にすれば欲しい答えが返ってくるかも? note LLM / imp 45
- LL.M.留学で持っていってよかったもの・いらなかったもの note LLM / imp 0
- 空白には、ルイス・キャロルを—— Mistral 7Bは、チョムスキーをどう読んだか note LLM / imp 35
- 【生成AIニュース+】『ChatGPT Business プレミアムシート』『Jalapeño』『Anima Turbo v1.1』『Ox Alpha』『ElevenLabs Composer』『10Eros-Max TURBO Hybrid Beta3』『MiniMax-H3-Longvideos』『MiniMax-H3-Fun-Controlnet-Union』『H3 Cinematic Multishot Coverage』『H3_Character_Sheet_Generator』他多数 note LLM / imp 30
- AIの文章に「透かし」を入れるということ——私たちは文章の血統書を必要としているのか note LLM / imp 45
- 3-7 学習データのバイアスと検閲の相関——「知らない」のか「知っていて隠す」のか note LLM / imp 50
- 『Anki』~なんでもAIに聞ける時代にあえて「覚える」ということ~ note LLM / imp 35
- MCPの新ロードマップ公開、今後はAIエージェント対応、HTTP通信への統一、アイデンティティ、よりよいデベロッパー体験などに注力 Publickey / imp 75
- Next.js 16.3正式版リリース。Turbopackのメモリ使用量が最大90%減、SSRが最大22%高速化、TypeScript 7による型チェック高速化など Publickey / imp 60
- AI活用率100%のQA組織をつくるまで LY Corp Tech Blog / imp 55
- LINEヤフーのAgent iを支えるAIエージェント基盤:「誰でも作れる」と「安全に動かせる」をどう両立したか LY Corp Tech Blog / imp 55
- Oktaから内製IdPへの認証基盤移行(第3回) DeNA Engineering / imp 50