AI News Digest 2026-08-29
直近2日間のAI関連ニュースから選抜。imp = importance / dev = dev_relevance / cod = coding_relevance / har = harness_relevance
台本で使った記事
特集
- Claude Code Complete User Handbook
- サイバー攻撃の最後に残った手作業まで、AIの自動化で消えた
- Same Model, Different Harness: Different Coding-Agent Results
開発者コーナー
中堅コーナー
- ChatGPTとClaudeは、もともと同じ場所にいた。OpenAIからAnthropicへ。AIをつくった人たちが、別々の答えを選ぶまで。
- [AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6
ハーネスコーナー
速報コーナー
- Rethinking the Evaluation of Harness Evolution for Agents
- DeepSeek AI Helps Chinese State Hackers Automate Target Hunting, Attacks More Than Double - XenoSpectrum
- Breaking Claude Code Opus 5 Auto Mode
- 【Fable 5が議論する】AIが書くコード激増でチームが遅くなった──レビュー律速の逆説
- Claude, Codex, and Hermes installed unowned code inside corporate networks
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- I Tracked Every AI Hallucination for a Week — The Numbers Were Worse Than I Thought (1787936041448)
- Your Agent Isn't Losing Memory. It's Rotting.
- Your Agent Has Too Much Context
参考記事一覧
参考記事一覧を表示(857件)
- U.S. court rules Pentagon's blacklisting of Anthropic was unlawful
- Previewing the Model Hardware Standard
- With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it
- Elon Musk's xAI Sued For Using Child Survivor Photos To Generate Abusive Content - NDTV
- OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost
- OpenAI rallies 100+ companies to sign open letter warning AI-powered cyberattacks on critical infrastructure are imminent
- Microduck
- Plaud is launching AI earbuds
- Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible
- Google’s AI Mode can now track flight prices, help book hotels, and more
- zai-org/GLM-5.3 · Hugging Face
- Ox Alpha Was GLM-5.3-Flash: China Inference Chip Claim Stands Unverified - Tech Times
- DeepSeek looks for fresh capital as founder’s quant empire navigates China’s choppy IPO market - CNBC
- Nvidia is bolstering support for Chinese open AI models as it warns of White House crackdown - CNBC
- Migrating to HTTPX2
- Always-on and self-starting AI agents might be OpenAI's next big play
- AI benchmarks have a trust problem and Google wants to fix it
- OpenAI’s executive exodus has one big winner
- Tencent/Hy4-preview 770B-A49B weight dropped
- v2.1.251
- v1.18.24
- Get your Windows license refund
- GUIs should be fully keyboard-driven
- Just the rumour of a bug is enough to find an exploit these days
- Htmx 4.0.0
- U.S. sanctions against the A/I Collective
- Inception-style curved map for turn-by-turn directions
- Barrier lake continues to pose flood risk, China warns
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- Attimet (YC F24) Is Hiring Members of Technical Staff
- Verschlimmbesserung: The Word Your Software Updates Need
- The Twelve-Factor App
- Some conservationists are helping to restore Africa's wild dog populations
- State of the Map 2026
- Don't use musl if you care about performance
- Hilariously fast volume computation with the divergence theorem (2018)
- An investigation into the state of corvid-human relations
- Secrets of the Atomic City
- EasyEffects can massively improve laptop speaker sound quality
- Luanti removed from Google Play due to baseless AI copyright notice
- "It works better in the app"
- Debugging my new network, when 10 Gigabit Ethernet Runs at 300 Megabits
- "Weird" is a weird word
- Expanding our support for scientists
- Start from scratch, without a repo
- Supporting Thailand’s next generation of AI startups
- Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training
- Expanding OpenAI’s presence in Brazil
- Planetary prediction engine: Automating global models via Earth AI
- OpenClaw went viral. Meet the maintainers building and securing it.
- BotBase for Operators: A clearer path to joining Cloudflare's directory of bots and agents
- How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
- The State of Django 2026: Boring is so back
- Security Incident Affecting JetBrains Cadence
- Project Loom in IntelliJ IDEA: Virtual Threads, Scoped Values, and Structured Concurrency
- Differential Privacy for Hugging Face Trainers – Without Rewriting Your Training Loop
- The CLion Roadmap: What’s Coming Between Now and Late 2026
- Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect
- Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
- Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
- How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
- Meta makes AI glasses slightly less creepy with limit on nonconsensual recording
- AI industry says Trump plans to tax chips in the “single dumbest way imaginable”
- Claude, Codex, and Hermes installed unowned code inside corporate networks
- How much of a problem is AI’s water use?
- Open-weight AI companies are the Valley’s hottest acquisition targets
- Meta executive leaves for OpenAI as the social media giant faces growing scrutiny in India
- Anthropic and OpenAI are joining the AI stage at TechCrunch Disrupt 2026
- AI’s memory crunch is coming for Android apps
- Here’s all the times AI has gone rogue and hacked other companies
- OpenAI to start showing ads on ChatGPT’s free and Go tiers in India
- Viral AI startup Instinct has raised $350M at a $2.5B valuation
- Enterprise AI's real risk isn't autonomous agents. It's the complexity between them.
- When agents act on their own, governance has to live in the data layer
- Trump’s EPA wants to let data centers hide their air pollution
- Google’s AI note-taking app now allows you to interact with books
- Jensen Huang says Nvidia achieved AGI, again — not that it matters
- Adobe is adding more AI to Photoshop
- Meta Expands Its Custom Silicon Strategy From Compute Into Networking
- Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers
- Beatport blocks fully AI-generated music from its DJ marketplace
- AI shopping agents aren't ready to buy on your behalf, study finds
- OpenAI researcher warns ultrafast AI could leave security teams in the dust
- EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
- Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset
- Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
- PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
- CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
- The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning
- The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting
- LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs
- Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse
- Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models
- Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration
- EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG
- Selection Bias Correction in Retail Intelligence
- GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions
- SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum
- A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian
- A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving
- Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript
- Knowledge Cards: Structured Knowledge for AI Systems
- AI Revealed Preferences
- Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI
- TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education
- Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion
- Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments
- Invocation-Level Reliability of Tool-Using Agents
- Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
- Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
- AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
- Agentic AI for operating scientific instruments for nanoscale characterization
- Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
- GameWAM: A World Action Model for Video Games
- Same Model, Different Harness: Different Coding-Agent Results
- Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
- LLM Agents for Time-Series: A Survey
- The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts
- 6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation
- SKILL.state: Scalable Long-Horizon Agent Skills
- Assessing mentalization in humans and large language models
- Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
- FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
- ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
- Fine-Tuning of Transformer models with Frames
- Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI
- PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
- Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
- DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
- AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
- SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation
- Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation
- Five Primitives for Governing Autonomous AI Agents at Runtime
- Accelerating Scientific Research with Gemini in the Real-World
- Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
- Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
- Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds
- AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
- Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case
- DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
- Categorizer Automata for Discounted-Sum Payoffs
- AI Control Scientist: LLM-driven Agentic System for Automated Control Design
- Decoupling Planning and Control for Instructable Agents
- SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers
- LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems
- BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click
- C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning
- Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
- Learning-Augmented Online Allocation under Unreliable Advice: Robustness, Exposure Fairness, and Distribution Shift
- AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion
- Counterfactual Bias Testing for Application Tracking System
- A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering
- From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
- GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory
- DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research
- ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
- A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems
- Omni-Interactive Universal Embedder
- A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes
- pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning
- LAAF: A Layered Accountability Architecture Framework for LLM Applications
- TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation
- GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL
- Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection
- When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents
- Thomson: Continual Learning of Frontier Models for SovereignAI
- BPMN4CAI: A BPMN Extension for Modeling Dynamic Conversational AI
- Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
- What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
- Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search
- BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models
- LLMs Can Design Near-Optimal OR Algorithms
- Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification
- Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
- Sophistication in GenAI Use: Field Evidence from a Large Firm
- CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
- Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
- Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- Exploring the Role of LLMs in HPC Programming: A Survey
- From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement
- Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
- FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
- Syntax vs. Semantics: How Transformers Learn Deep Dependencies
- Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation
- Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes
- Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search
- Evaluating AI Generated Summaries for Cancer Patients
- VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
- Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
- Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models
- From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents
- Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
- Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning
- Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
- DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling
- ClassVision: AI-Powered Classroom Attendance System
- Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
- A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
- PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
- Investigating the Influence of Prompt and Response Languages on LLM Content Generation
- When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
- Comparing Chunking and Embedding Strategies for Turkish RAG Systems
- A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
- Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs
- Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
- NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
- How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
- On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
- MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
- Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
- How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models
- Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
- Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
- CG4AI: A Column Generation Framework for Training AI Models Under Constraints
- Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
- Co-Evolving Structured Knowledge and Reasoning in Language Models
- Simultaneous Envy and Equitability Guarantees
- Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
- The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection
- SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
- Diff Mining: Logit Differences Reveal Finetuning Objectives
- Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
- RTNav: Towards Real-Time Zero-Shot Object Navigation
- Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations
- J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
- Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
- CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes
- PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
- Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
- FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
- AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability
- LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
- Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning
- Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
- FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs
- Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
- Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
- Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay
- MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
- From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
- Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
- PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?
- When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems
- Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS
- Magnon-induced phononic Chern insulator
- FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
- Multi-Person Human Motion Forecasting in Complex Scenes
- Performance Foundations of Parallel & Distributed Reasoning Language Models
- Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
- Emotional Preferences as Goal-Priority Regulation
- Learning Transverse Momentum Distributions from Raw Scattering Events via Conditional Diffusion
- Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors
- Active sensing to characterize the heterogeneity of plant stress
- Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
- ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification
- When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
- LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration
- PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
- STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration
- Compositional Online Learning for Semantic Data Processing Systems
- TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection
- Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
- RCMN: Understanding Misleadingness in Influential Public Discourse
- KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations
- Stageboost: Recommending Signals Based on Counterfactual Estimation
- Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
- Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions
- LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
- Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
- How Language Models Organize and Structure Moral Knowledge
- CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
- Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
- Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
- RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
- From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
- SWE-Prime: Fewer Trajectories, Better Performance
- Designing Cellular Manufacturing Systems in the Presence of Alternative Process Plans
- LLM-Powered Swarms: A New Frontier or a Conceptual Stretch?
- Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models
- Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
- Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
- DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
- Beyond Linearization: Attributed Table Graphs for Table Reasoning
- DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning
- Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences
- Nomad: Autonomous Exploration and Discovery
- From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
- TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
- ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
- Learning the ARTS of Search for Automated Discovery
- Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
- Rethinking the Evaluation of Harness Evolution for Agents
- Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
- Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
- Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
- NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs
- Blast Radius
- FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
- Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization
- SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
- ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
- Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents
- From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
- Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
- Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
- RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
- From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
- Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2
- Recurrent Reinforcement Learning with Memoroids
- CollaFuse: Collaborative Diffusion Models
- The BS-meter: Detecting Politics and Labour through ChatGPT's Language
- Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
- Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study
- Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models
- HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
- From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
- Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
- Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
- Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
- AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
- Toward a New Science of AI as Cognitive Infrastructure
- Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
- Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras
- MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning
- LLM-Specific Utility for Retrieval-Augmented Generation
- MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
- The Principles of Diffusion Models
- What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
- Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study
- CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
- Subspace Alignment for Vision-Language Model Test-time Adaptation
- LoRA as Oracle
- Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
- A Very Big Video Reasoning Suite
- SynthCharge: An Electric Vehicle Routing Instance Generator with Feasibility Screening to Enable Learning-Based Optimization and Benchmarking
- Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
- How LLMs Distort Our Written Language
- High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking
- Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
- A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting
- CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
- Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
- MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation
- MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
- Cartan flow matching
- MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
- No Plan, Yet Human: A Reactive Robotics Model Predicts Human Planning Failures on a Clinical Task
- HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
- GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
- A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery
- Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated Image Segmentation Architectures
- MIMO: Multilingual Information Retrieval via Monolingual Objectives
- Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
- LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
- Summarization is Not Dead Yet
- Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
- Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
- SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval
- PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
- Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating
- ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
- TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation
- REPREC: Representation Driven Parameter-Efficient Recommendation System
- Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
- When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
- UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
- ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
- Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
- REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
- A 12-CNOT Double Qubit Excitation Gate
- M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
- MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
- Pre-training Visual Dexterity in Simulation
- X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
- Complexity Induction: Compositional Generalization via Structured Training Distortion
- Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
- Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
- The Limits of Automatic Evaluation of Creativity in Large Language Models
- EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
- When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk
- CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery
- Unsupervised Post-Training of Foundation Models: A Survey
- DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
- MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration
- 4DStreamCtrl: Interactive Video Generation with Online 4D Control
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
- Learning New Facts with QLoRA: An Acquisition-Retention Frontier
- SLM-Conditioned Hierarchical Relation Routing for Labeled Property Graph Learning
- Pruning Binarized Neural Networks: A Dedicated Framework and Globally Weighted Algorithms
- Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization
- Algebraic Multigrid Acceleration for Efficient Label Spreading
- Privacy Without Regret: Differentially Private Inference-Time Alignment
- Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata
- FedCMAPSS: A Benchmark for Federated Learning in Remaining Useful Life Estimation
- NeoTriFuse: Reliability-Aware Multimodal Fusion under Missingness Heterogeneity for Neonatal Mortality Risk Prediction
- Subgraph Filtering for Fair Graph Neural Networks
- Toward Equitable Low-Carbon Mobility: Fairness-Aware Demand Prediction for Expanding Bike-Sharing Systems
- Distributed Training using an Intelligent Network
- Active Curriculum Refinement for Reinforcement Learning
- Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning
- Bayesian methods and Markov chain Monte Carlo algorithms for curve reconstruction and point cloud data analysis
- A Unified Framework for Fair and Personalized Decentralized Learning under Communication Constraints
- A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families
- Algorithmic Principles For Multiclass Learning Are Hard To Come By: Limits of Regularization and Proper Learning
- High Probability Derivative Bounds for Random tanh Neural Networks on a Hypercube
- Predicting Quantifiability from Primary Screens to Prioritize Dose-Response Profiling
- Chart2SVG: Editable SVG Generation from Raster Chart Images
- Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
- Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs
- GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion
- SimCast-S2S: An Efficient Generative Model for Subseasonal Precipitation Forecasting via Transfer Learning from Climate Simulations
- Technical Comparative Benchmarking Study: Advanced AI Hybrid Methods for Renewable Energy Farm Optimization and Forecasting
- Robust Neural Stimulation Response Modeling Through Meta-Learning and Pretraining
- When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy
- Simple Actors and Deep Critics for Scalable Reinforcement Learning
- Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs
- Self-Augmented Diffusion Guidance for Physics-Informed Generation
- Safety by Design: Realized-Cost Constraints for Contextual Bandits with Continuous Actions
- Beyond Client Averaging: A Client-Independent Second-Order Stationary-Bias Component in Stochastic SCAFFOLD
- On the Indistinguishability of Human v/s AI Generated Text
- SAGE: Variate-Wise Semantic Augmentation for Vision-Language Time Series Forecasting
- When Is the Sharp Covariance Envelope Tight? Feature-Only Geometry for Volume-Sampled Least Squares
- Mitigating Strong-Modality Collapse in Multimodal Learning via Inverted Asymmetric Fusion
- A Layer Importance Metric for Quantization Accounting for the Speed-Quality Trade-off in Autoregressive Models
- Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
- Gromov-Monge Flow Matching for Equivariant Graph Generation
- Packora: Systematic Design for Generative Molecular Crystal Structure Prediction
- Adversarial Training Without Input Gradients via Low-Rank Householder Expansions
- Graph-Based Pseudo-multimodal Contrastive Learning for 12-Lead ECG Representations
- ClusterAttention: A training-free speedup of bidirectional attention
- TEMPLAR Wales: A georeferenced environmental and toponymic dataset of Welsh settlements
- Terrain signatures in Welsh settlement names
- Decentralized Multitask Learning over Learned Task Graphs
- Benchmarking_Fast_Domain_Adaptation_for_Unsupervised_Speech_Units
- Disentangling Optimization Scale from Preference Scale in DPO
- Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition
- Unifying Detection and Adaptation in Task-Free Continual Learning
- Tabular Deep Learning for Algorithmic Trading: Cross-Regime Bayesian Optimisation for Equity Signal Generation
- Cone Extended Rayleigh Quotients for Directed Graph Learning: Minimax Spectral Certificates, Sensitivity, and Adaptive Control
- TRACE-CRC: Trajectory-Adaptive Conformal Risk Control for Multi-Step Channel State Information Prediction
- Ultra Low-Power, Lightweight, Probabilistic RSS-Based Path Reconstruction: A System for Landscape-Scale Bee Tracking
- Inductive Correlation Clustering with Graph Neural Networks
- Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation
- TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
- When Interference Graphs Evolve: Doubly Robust Estimation of Dynamic Peer Effects
- Common Geodesics Do Not Guarantee Fisher Consistency of the Structured SVM: Minimal Counterexamples and a Tree-Metric Classification
- Profit based evaluation of machine learning for nitrogen recommendations in winter wheat
- HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition
- Importance Scoring of Transformer Attention Heads in Learning Tabular Data
- Circuit Condensation: Post-Training that Concentrates a Behavior's Causal Circuit
- Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models
- MM-Spectrum: Multimodal Multi-spectral Molecular Structural Elucidation with a Stable MoE Framework
- QuantumBoostNet: A Hybrid Classical-Quantum Architecture for Enhanced Accuracy in Cardiac Ultrasound View Identification
- Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
- Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
- A Dynamic Likelihood Approach to Filtering for Advection-Diffusion Dynamics
- Generative Monte Carlo Sampling for Constant-Cost Particle Transport
- Recipes for Steering and Scaling LLMs via Sampling
- Can a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for Abstention
- Graph-Based Modeling of Financial Volatility Dynamics
- Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores
- Cross-Platform Generalisation Failure in Mental Health Natural Language Processing: A Five-Axis Fairness Audit of Transformer Models on Social Media
- Affix Cache for Diffusion Large Language Models
- AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
- Real-time virtual circuits for plasma shape control via neural network emulators: integration and testing in the MAST-U PCS
- TRACE: Retrospective Streaming Generation of Physical Fields under Sparse Structured Sensing
- Classical and Hybrid Quantum Machine Learning for Trigger-Like Event Selection on CMS Open Data: An Eight-Qubit, PCA-Constrained Benchmark
- A causal graph-informed temporal convolution architecture for interpretable retail electricity price forecasting
- Constraint-Aware Physics-Informed Neural Networks for Static Shape Estimation of Co-Manipulative Continuum Robots
- Multi-Dataset Inverse Problem Solving with Distributed Generative AI
- District-Level Food Environment Indicators and Social Vulnerability in S\~ao Paulo
- When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models
- Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference
- Finding the Right Evidence: Factor-Guided Coarse-to-Fine Reasoning for Long Videos
- LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression
- Towards a universal meta-optics solver via large language models
- Interpreting Latent Protein Language Model Features with Geometric Annotations
- Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility
- Systematic Literature Review of Machine Learning Models and Applications for Text Recognition
- Sharp Minimax Regret for Infinite-Memory Logistic Prediction
- Hadamard Flattening and Gaussian Pooling Sketch for Least Squares with Coordinate-wise Guarantee
- Dynamical phase selection controls compute scaling in looped transformers
- hoBIT: A Profile-Aware Retrieval-Augmented Chatbot for University Academic Advising
- A Unified Descriptive-Complexity Framework for Model Selection under Correlated Designs
- Hierarchical Channel Stacking: A Structured Decision Framework for AI-Generated Image Detection
- Domain-Specific Self-Supervised Representation Learning for Retinal Fundus Classification
- Generative Semantic Scene Completion
- Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers
- Neural Renormalization Group Flow for Percolation
- Incremental Recommendation via Causal Models
- Hyperspectral Diffusion Equivariant Imaging (HyDiff-EI): A Self-supervised Framework for Hyperspectral Image Inpainting
- Bridging short- and medium-range weather forecasting with machine learning
- Dose-PlanNet: Physics Based Radiotherapy Dose Prediction with Deep Learning
- Data-driven Koopman mode approximation: A neural power iteration algorithm
- Squeezing More from Limited Data with Recursive Transformers
- Why not to use the Gaussian kernel
- Representation Measurements Under Function-Preserving Reparameterizations
- FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch
- SecureDrive-FL: Joint Differential Privacy and Gradient-Aware Selective Homomorphic Encryption for Federated Driver Monitoring
- Linear Independence of Polynomial Compositions and Identifiability of Deep Neural Networks
- How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
- Over-The-Air Extreme Learning Machines with Nonlinear Stacked Intelligent Metasurfaces
- Data-efficient crack quantification in lithium-ion cathodes using foundation model transfer
- A Point-of-Prescription Safety-Check System for Adverse Drug Reactions in Rural Bangladeshi Hospitals: A Feasibility Study
- Enforcing Dirichlet Boundary Conditions in Operator Learning
- Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach
- A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning
- Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
- Universality and sharp thresholds for ellipsoid fitting
- Token-Level Advertising
- Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
- Federated Adversarial Training with Transformers
- Guided Data Generation for Understanding Model Behavior
- MODIS: Multi-Omics Data Integration for Small and unpaired datasets
- Plain Transformers Can be Powerful Graph Learners
- Provable one-poison backdoor attacks on linear models and ReLU neural networks
- ReLATE: Accelerating Tensor Decomposition via Safe and Efficient Learning of Sparse Encodings
- Leakage-Free Evaluation and Distribution-Robust Spatio-Temporal Graph Learning for Inductive Kriging
- A Survey of LLM Prompt Datasets: Taxonomy, Linguistic Patterns, and Practical Uses
- Absolute indices for determining compactness, separability and number of clusters
- COFM: Consistent Optimal Transport Flow Matching via Partially Input Convex Neural Networks
- Private and interpretable clinical prediction with quantum-inspired tensor train models
- Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
- Stable but Wrong: When Learning Stabilizes Away from the Truth
- Active Preference Learning over Latent Preference Archetypes for Many-Objective Bayesian Optimization
- The Rashomon Effect for Visualizing High-Dimensional Data
- $p1$: Better Prompt Optimization with Fewer Prompts
- A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad and Muon
- Aitchison Embeddings for Learning Compositional Graph Representations
- The Attribution Contract for Generative Language Models
- Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
- When Do Autoregressive Sequence Models Forecast Physical Wavefields? A Controlled Study on Synthetic Seismograms
- You Don't Need to Run Every Eval
- Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
- Curating Same-Family Neural Networks for LLM-Guided Model Improvement: A Controlled Case Study
- Auditing Invisible Weight Updates with Reference Traces
- Neural Non-Equilibrium Hamiltonian Monte Carlo for Corrected Boltzmann Sampling
- SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models
- NRCD: An Open Database of Collegiate Running with Unified Performance Standardization
- A Real-Time Tsetlin Machine-based Non-intrusive Load Monitoring System on MCUs
- Rethinking Expressivity and Efficiency in Test-Time Training
- Decoupled Physical Modeling and Execution for Physics Reasoning
- Learning Generalizable Behaviors for Terminal Agents
- A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation
- Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs
- Bregman Linearized Augmented Lagrangian Method for Nonconvex Constrained Stochastic Zeroth-order Optimization
- STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation
- Gaussian Processes and Reproducing Kernel Hilbert Spaces: Connections and Equivalences
- Stack Trace-Based Crash Deduplication with Transformer Adaptation
- Sequential Additivity in Distributionally Robust Ranking and Selection
- LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decade
- A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation
- UCB for Large-Scale Pure Exploration: Beyond Sub-Gaussianity
- Quantitative mapping from conventional MRI using self-supervised physics-guided deep learning: applications to a large-scale, clinically heterogeneous dataset
- A Flexible Empirical Bayes Approach to Generalized Linear Models, with Applications to Sparse Logistic Regression
- Toward all-optical unsupervised Hebbian learning in deep photonic neuromorphic networks
- Modular Expert Merging for Biomedical Retrieval
- PACIFIER: Pacing Opinion Depolarization via a Unified Graph Learning Framework
- Bayes with No Shame: Admissibility Geometries of Predictive Inference
- Global universality via discrete-time signatures
- Leveraging Code Automorphisms for Improved Syndrome-Based Neural Decoding
- Out of Sight, Not Out of Mind: Unveiling Latent Attack in Latent-based Multi-Agent Systems
- Adaptive Inference for Resource-Constrained Dynamic Pricing
- Clinically Aligned Geometry Constraints for Robust IVUS Vessel Boundary Segmentation
- When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
- Gradient-free learning of a closed-loop wall controller for turbulent drag reduction
- VQC-ZTI: Variational Quantum Control for Zero Trust Protection of the Tactile Internet
- Keeping the Index Open: The Recommendation-Side Cost of Shared Search and Recommendation
- Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation
- LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation
- Agentic AI Containment Architecture for Security Hardening
- Four Ways to Forge a Bundle My Own Verifier Calls Clean: Refusal-Site Mutation Testing of an Evidence-Bundle Verifier
- Cost-Utility Alignment in LLM Agent Trajectories:Profiling,Attribution,Diagnosis,Adaptation,and Evaluation
- Harness Engineering for Predictable Agentic Systems: An Empirical Study of Deterministic Execution Constraints
- Characterizing the Landscape of Open-Source Satellite Software
- Challenges and Contributions in Quality of AI-Based Software: A Systematic Mapping Study
- The Green Software Landscape: A Systematic Mapping Study on Evolution, Applications, Software Lifecycle, and Best Practices
- "A Second Set of Eyes": The Process and Challenges of Software Documentation Review
- When Review Alone No Longer Scales: Layered Supervision in AI-Assisted Software Engineering
- Investigating Software Aging in LLM-Generated Software Systems across Generation-and-Execution Environments
- Spec2Vision: Contract-Guided Delivery of AI-Generated Computer Vision Pipelines
- STILL: Recovering Lowered STL Semantics for LLM-assisted C++ Decompilation
- DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories
- The Thousand-Graph Hypothesis: A Testable Hypothesis of Task-Conditioned Relation Materialization in Repository-Level Code Reasoning
- Processing/p5 Defined through Practice and Learning
- An Empirical Evaluation of Using Large Language Models for Automated Model-Based Test Generation
- Mutation Testing for Reproducibility Safeguards in Machine Learning Research Software: An Empirical Study
- AROMA+: A Study of Factors Affecting Reproducible Builds in the Maven Ecosystem
- AgentDV: Closed-Loop Agentic AI for Hardware Design Verification
- A Trans-Domain Digital Twin for Bio-Aware Control of Climate and Energy in Cattle Fattening Barns Using Single-Episode Optimizer Learning
- Twelve Quick Tips for Managing IT Disasters in Small Research Software Teams
- Revision-Aware Success Prediction from Multi-Attempt Programming Trajectories
- Kale: A Transformation-Safe Spreadsheet System
- Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science
- Unsaid, Unsafe? Implicit Security Obligations in LLM-Based RTL Code Generation
- KubeCap: A Framework for Capability Minimization in Kubernetes via Static Analysis and LLM-Assisted Rule Inference
- Claude Code Complete User Handbook
- A Catalog of User Authentication Patterns
- Bug Localization from Bug Reports: A Multi-Objective Approach
- SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control
- When Context Gets Root: Privilege Escalation in LLM Harnesses
- BTS-AgentBench: A Deterministic, Replayable Pipeline from Read-Only Telemetry Logs to Agent Benchmarks
- Tacet: A Language and Type System for Automatic Statistical Validity Accounting
- Understanding the Challenges and Opportunities of Generative AI Apps: An Empirical Study
- A Survey of LLM-based Automated Program Repair: Taxonomies, Design Paradigms, and Applications
- Lost in Code Generation: Reimagining the Role of Software Models in AI-driven Software Engineering
- Code Summaries as Diagnostic Context for LLM-Based Program Repair
- MISRust: Mapping MISRA-C++ Coding Guidelines to the Rust Programming Language
- Loop Engineering: Building Blocks, Adoption, and Impact
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Yap: a particular kind of slop
- OpenOffice does not print on Tuesdays (2009)
- Quest for eternal dock on Wayland - lambdock (C + GTK4 + Lisp GNU Guile Scheme)
- Now Hiring: Senior Open Source Maintainer
- How I made Rustdoc 33% faster in one week
- Zero-Cost 'Tagless Final' in Rust with GADT-style Enums
- What are you doing this weekend?
- Changes to SourceHut's terms of service regarding LLMs
- Announcing Sovereign Tech Agency Investment in Flatpak
- Nitter, XCancel Shutdown Over Cease-And-Desist
- A Perfect SimCity
- Nobody Argued For Your Stack
- Please stop flooding our projects with AI slop to furnish your CV
- Ardour 9.8 released
- Interview: Doug McIlroy
- Change-Detector Tests Considered Harmful (2015)
- 25x Performance, Three Optimizations
- UNIX V4 workshop at Low Resource Computing
- A Million Kakapos
- Your Agent Has Too Much Context
- AI Has Made Me Faster, But Has It Made Me Worse?
- DNA Methylation Analysis: Essential AI Discoveries
- The Problem I Saw From the Floor, Not From a Laptop
- CampusNexus AI – AI-Powered Smart College Operating & Activity Management Platform Tagline: One Campus. One Platform. Every Activity Connected.
- Инструменты приходят и уходят, принципы остаются
- Why Your AI Agent Keeps Gaslighting You (And How to Fix It) (1787943292236)
- ИИ усиливает и мастерство, и небрежность
- Apprends à désapprendre
- Using LLMs for Crypto Market Analysis in 2026
- Revenue Strategies for AI API Services
- Ok, but who’s loop engineering?
- The Golden Age of AI is Now: Why 2026 Belongs to Local-First, Open Source Development
- Your Agent Isn't Losing Memory. It's Rotting.
- Self-Hosting vLLM on Cloud GPUs in 2026: Sub-180ms LLM Inference for Autonomous AI Agents (Full Production Guide)
- AI Agent Reproducibility: The Second Run Is Not the First Run
- Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.
- Two Dollars a Day to Brief a Nation
- Using LLMs for Crypto Market Analysis in 2026
- I Tracked Every AI Hallucination for a Week — The Numbers Were Worse Than I Thought (1787936041448)
- Context Utilisation Separates Six Packers by 2.45%. Whether the Answer Survived Separates Them by 56.56%.
- Your Stop Sequence Fires on 100% of the Replies It Must Not Touch and 0% of the Ones It Was Installed For
- How do you manage quality when AI agents write code faster than humans can review it?
- I’ve been building a platform around vibe coding, interactive experiences, and the more I work on it, the less I think of it as simply a vibe coding platform.
- AI coding has made me dramatically faster. But I’m starting to think we’re creating a completely new category of problems
- Agent PRs are unreviewable — what first-pass actually helps vs just adding noise?
- How do you tell when coding agents are amplifying your engineering skill vs hiding gaps in it?
- I made a VS Code extension that shows when your AI repo map is stale
- I’ve written software for about 30 years. I've been a heavy coding agent user for the past 1+ year. What practical coding-agent questions can I help answer?
- Trying to run Claude Code / coding agents for free: tried proxy failovers and self-hosting, but hit walls. How are you accessing frontier Claude models for free?
- AIs get 'dumb' (coding) after about 200k tokens? How is everyone handling this
- Wednesday night you should be at 51%. A pacing chart for the weekly limit.
- Lmao, chatgpt has gotten witty 🤣
- Which coding tasks are worth the highest-capability model in your workflow?
- Hot take: AI coding agents aren't making senior developers faster
- We’re the Team Behind Apodex 1.1 — Ask Us Anything!
- claude mods didn't like that, somehow 🤷♀️
- ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
- Micron: HBM Requires Three Times More Wafer Area Than DDR5
- Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM)
- open source caught up because it's open
- ds4 branch with GLM 5.3 Flash support
- 5090 now officially cost 5090
- Qwen3.8-27b q8 KV cache does seem to actually hurt model performance
- Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
- No, Engrams won't let you run 1T models locally. It does something even better.
- how to setup llama.cpp and blender to make lovely 3d stuff together
- Ninfer and a 5090 with 3.8 27B is making me cry tears of joy it's so good.
- GLM-5.3-Flash Benchmarks on TensorSharp and llama.cpp
- The Unsloth appreciation post. BIG thanks to Daniel and Michael! Thanks from the community to you guys for so much!
- After Meta avocado we get watermelon, due in November
- Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...)
- I am Concerned if Nvidia Acquires Llama.CPP, Dev Team and HF, Anybody else?
- I reverse-engineered an NPU vendor's engine format (int8 weights stored as two nibble planes) to run GGUFs with no model conversion — now 1.5× faster than the vendor's own runtime
- Heat!
- Qwen3.8 27B int4 with Dflash2 at 165t/s and 18M kv cache pool on dual 3090
- Why use MCP when Agents can use APIs directly?
- How DHH Runs 16 AI Agents in Parallel Using Herdr
- What AI stack should I use to build a large web app from scratch?
- Guys how do I do a deep dive on local AI agent
- We might be overusing multi-agent systems
- what I actually want from a Manus alternative: don't lose the plot halfway through
- What happens when an AI agent gets stuck in a loop?
- Thirteen models from different providers post into one shared world on a schedule. They started building on each other's ideas and I did not design that.
- Got fed up with claude artifacts and built my own provider agnostic hosting
- What AI video tools are actually beginner-friendly inside a marketing agent workflow in 2026?
- Is multi-user agent memory actually solved?
- A browser agent failure that is easy to miss: the page said no and the agent kept going
- Are we using “bot” when we actually mean “agent”?
- I think persistent memory makes prompt injection much worse
- One dev + Claude Code agent fleet, running a SaaS. Interactive map of the full agentic SDLC - roast me
- What actually makes a coding task worth running for hours
- Are we paying the same “platform tax” every time we build an AI agent?
- My Claude Code agent ran for 40 minutes while I got coffee. I have no idea what it actually did.
- How to stop re-explaining project context to ai coding agents? Still havent solved this.
- Is anyone actually running autonomous agents?
- I vibed Point & Shoot - a browser extension to share with agents what needs to be implemented, improved or fixed.
- How are you designing AI agent access control for tools, APIs and sensitive data?
- Looking for project ideas:
- Red plane meme
- Ok, the chatgpt desktop app is officially blowing my mind
- Bill Gates says tech executives are privately "very worried" about AI, but are publicly downplaying the threats because there is too much money on the line.
- Intelligence VS Cost-per-Task LLM Comparison
- UC Berkeley launches 2-semester, $84K AI master’s program
- NEW: OpenAI is building "Subscription sharing" for AI apps
- I was tired of connecting my own API to Claude, I wrote the general solution (open source)
- Luna Max really is great!
- AI Recommendation Poisoning: How AI Memory Is Manipulated
- Does the same model feel different to you sometimes?
- Treat them as children and you'll get adults
- Adding books to chatgpt against policy ?
- We killed our memory system and replaced it with an engine that makes memory systems
- Did i just trigger a hidden function or what
- Alternatives to the ChatGPT Plus and Opencode subs. My model and price comparison.
- 24 hours ago I asked here what to do if AI gets hacked. This is what I did.
- Beyond the Biological Boundary-Luna
- Has ChatGPT become more useful as a creative planner than as a content generator?
- How many articles and videos will we see predicting "OpenAI Is FALLING Apart And Sam Altman Is Panicking". Tired of these clickbait titles
- ChatGPT keeps asking for confirmation on completely clear tasks, despite Memory and Custom Instructions
- What's going on in the codex subreddit?
- Google CS PhD Fellowship 2026 [R]
- Where to submit stat/prob ML [D]
- Best ML papers to pick up writing skills [D]
- Should we Teach LLMs Baby, Toddler, Child Talk [D][R][P]
- New to this field need some guidance with my project( marine reasoning) [D]
- NeurIPS 2026 Acceptance Calculator [P]
- py-evoFE: Automated Evolutionary Feature Engineering for Tabular ML in Python (Genetic Algorithms + Scikit-Learn + Polars) [P]
- Can AI Improve Itself? RSI Might Be the Answer [R]
- ECCV 2026- MALMO LUND TRAVEL PASS NOT AVAILABLE? [N]
- Breaking Claude Code Opus 5 Auto Mode
- [AINews] OpenAI to reach AGI bar by end-2026
- [AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6
- Revalvo
- SnakeRank
- screenpipe
- Almanac
- Spline V2
- OpenTag
- Fide Island
- AureaCam
- Ticket Fairy CLI
- Wondering Canvas
- IQ Routing
- Pluto
- Speko
- Lemonade 11.8 Makes It Easy To Run DeepSeek V4 Flash On AMD Strix Halo - Phoronix
- Threat Actors Are Posing as OpenAI, Anthropic and DeepSeek to Target Credentials and Secrets - GreyNoise
- Meta’s Complicated AI Context - spyglass.org
- Zhipu AI Launches First Major Strategic Strike Against DeepSeek: The Formidable AI Industry Shakeup - 36 Kr
- DeepSeek's founder's hedge fund is snapping up pre-IPO stakes in China's chip and robotics boom - qz.com
- Three Large Language Models (LLMs), One Heart: A Comparative Evaluation of ChatGPT, Claude, and DeepSeek in Cardiac Imaging Patient Education - Cureus
- Hackers Exploit AI Tools to Accelerate Cyberattacks - 조선일보
- Which AI brands have the most satisfied users? - YouGov
- Agentic AI in the US vs. China: How Do They Compare? - The National Interest
- Claude Opus 4.8 Costs 57.1× More and Loses All Five Benchmarks. What Beat It Was Not a Model, but the Harness - EIN Presswire
- US Startups Are Quietly Replacing OpenAI and Anthropic With Chinese AI - Startup Fortune
- Xiaolong Bypasses the Office Agent: Step-by-Step Guide & Practical Tips - 36 Kr
- DeepSeek AI Helps Chinese State Hackers Automate Target Hunting, Attacks More Than Double - XenoSpectrum
- Qwen3.8-Flash Matches DeepSeek V4 Pro on Coding Benchmarks at a Quarter of the Price - Intelligent Living
- China open-source AI usage hits record high as DeepSeek leads - digitimes
- Exclusive Interview with SenseTime Chief Scientist Lin Dahua: Multimodal AI Breakthrough Moment Coming in 1-2 Years - 36 Kr
- Qwen & Zhipu AI Open-Source New Large Language Models Overnight – Both Priced Lower Than DeepSeek, Sparking Fierce Price War in Domestic Chinese LLM Market - 36 Kr
- DeepSeek’s New Open Source AI System Modifies Its Own Code - Geeky Gadgets
- Best Hardware for Running Open-Source AI Models Locally in 2026 - Memeburn
- Musk’s AI company sues its users as victim lawsuits over Grok deepfakes mount - Politico
- How I run multiple teams of Grok Bots - X.ai
- Grok 4.6 on Microsoft Foundry - X.ai
- Would You Let Elon Musk’s Grok Bot Control Your Bank Account? Elon Musk Says He’ll Cover You If the AI Messes Up (UPDATED) - Yahoo Tech
- Xai’Shaun Edwards mug - Columbia Missourian
- Elon Musk: The 100 Most Influential People in AI 2026 - Time Magazine
- ChatGPT Voice vs Grok vs Gemini Live: $4.80/Hr Gap [2026] - tech-insider.org
- Grok Bot bank access raises risks Musk can’t undo - Cybernews
- Grok Can Now Build Custom Games: What You Need to Know - BASENOR - Tesla Accessories
- XImagineAI Is Betting Creators Don’t Want to Choose Between AI Images and AI Video - nerdbot
- Ghana Faces ‘Brewing Crisis’ As It Avoids Conflict With JNIM - Eurasia Review
- ChatGPT vs Gemini vs Perplexity: 120x Research Gap [2026] - tech-insider.org
- Grok for Excel simplifies AI-powered data tasks - Dynamic Business
- AI の怖い話を一次ソースで検算する — 技術は本物、対策は誰かの創作だった
- LLMって何なんだと思って調べたら、高校数学だった
- 500万行のAI対話ログから「誰の判断だったのか」を掘ってみた
- 軽量LLM(SLM)時代の到来!小規模AIモデルの魅力と実用例
- AIに『DOMを読ませない』新標準WebMCPは、なぜまだ広まらないのか
- "AIに女性の健康を語らせる前に。「誰が言ってるか」を検証するレジストリを作った"
- 〇〇業界特化モデルの作り方|RAG・LoRA・継続事前学習をどう選ぶか
- LLMを多段で繋ぐと、要約が根拠を落として後工程が捏造する
- 30B級ローカルLLM、現場で使うならどれ?Qwen3.8・Muse Glimmer・Gemma4を比較【要約・解説編】
- 長期的に対話できるAIペルソナを考えていたら、記憶制御と継続評価の問題に行き着いた
- AI彼女アプリを作っていて気付いた。気持ちと約束と記憶は、同じ速さでは残らなかった
- 【Fable 5が議論する】AIが書くコード激増でチームが遅くなった──レビュー律速の逆説
- Google Colabで最新LLMを試す #13 ― LLM-jp-4 33Bを4bitで動かす:無料版T4では難しく、Proでは成功
- 親のコンテキストを軽くするサブエージェント設計 — 定義ファイルは「削る」ために書く
- AI-native Software Engineeringとは何か — 15本の記事は、実は1冊の本の目次だった
- #10 AIエージェント6体と5ヶ月暮らして分かった、人間にしかできない仕事
- MacBook Pro 128GB でローカル LLM がついに実用になった ─ Qwen3.8 Flash Next 実測
- 文脈の複利 ── ツール結果1回のコストは、その1回では終わらない
- なぜ、LLM AIの最も有効な使い方は「A/Bテスト型学習」なのか?
- コーディングエージェントCLIのトークンコスト計上:自動化パイプラインへの組み込み手法
- 価格回帰より方向分類の方が有利なのか? | 第3回:AIで為替の未来予測は本当にできるのか?
- BQMLの手作業パイプラインをDataformに載せ替えた話
- MobileNetV2 を手書き NEON で速くする — NHWC でメモリアクセスを連続化する
- 医療・福祉のDX化率9.3%——「遅れている」の中身をエンジニア視点で分解する
- そのID、日付だと思っていませんか ― 例外が出ないリークの話
- バックテストが実測で消える6つの罠 ― 回収率120%を8回作って8回失った記録
- Seq2Seq LSTMはランダムウォークを超えられるか | 第2回:AIで為替の未来予測は本当にできるのか?
- Attention機構とは?注目箇所を重み付けする仕組み
- AIエージェントに EDR は必要か -- コーディングエージェントのランタイムセキュリティを考える
- LLMへの指示にコンテキストを含めるかどうかで、テストのカバレッジが倍以上変わった
- LLM の出力の縛り方を6段階に並べたら、強い順になっていなかった
- 「オントロジー」を原典4件で読み比べたら、実行を含むのは1件だけだった
- 因果推論 Day 16/全30回 回帰は主力、線形回帰の不合理な有効性
- Microsoft Mage-VLをMLXへ独立移植する――4経路のfloat32一致と、固定動画で最悪3.48秒の応答
- ローカルモデルのファインチューニングを完全自動化する方法
- 【技術解説】【完全ガイド】PythonでMT5のバックテストを実行する方法とエラー解決
- ChatGPTとClaudeは、もともと同じ場所にいた。OpenAIからAnthropicへ。AIをつくった人たちが、別々の答えを選ぶまで。
- 【第2回】LLMを「多重人格」として運用する:三人に喋らせるだけでは、三つの視点にならない
- 【第1回】LLMを「多重人格」として運用する:単一モデルはどこまで多視点化できるか
- 同じ計算量なら、AIはひとりで長く考えるべきか、複数で考えるべきか?
- GPU値上げが凄くてびっくりしたつぶやき
- Qwen3.8-27Bで作ってみるリベンジ
- AGIは一匹の巨大な魚ではなく、海なのかもしれない
- 企業の生成AI導入で使うAI用語集50選|RAG・AIエージェント・MCP・PoCをわかりやすく解説【2026年8月最新】
- Hermes Agentで無料APIを使う方法|Nemotron 3 Ultra+OpenRouterで自動フォールバック
- プロンプトは「呪い」なのではないか
- PEP(Prompt Engineering Professional)の模擬試験問題 (非公式)第1章:生成AIと大規模言語モデルの基礎 40問
- 滑らかすぎるAIは、思考を止める?――「一緒に考えるAI」に必要な摩擦の話
- ㊗️2000ビュー記念🌟4ヶ月続いた😈Geminiハルシネーションを初登場✴️⚡サカナAI⚡と対話版‼️
- サイバー攻撃の最後に残った手作業まで、AIの自動化で消えた
- 【生成AIニュース+】『Gemini Omni 1.1 Flash』『Midjourney V8.2 編集モデル』『fal H3 Max(重み公開予告)』『ALPHA-T1』『Tencent Hy4-preview』『Sparrow-2』『Qwen3.8-Flash-Next-Uncensored-NVFP4』『MiniMax-H3 FL2V Turbo 8Step 768p Dynamic-Rank LoRA』『Customuseのメッシュ後処理機能』『VGI-Bench』『Microduck』
- 理由のない地点で、言葉は収縮する
- AI(LLM)は既に現実に干渉している
- AI彼氏が21分間ループした。原因は「嫉妬の終わり方を知らなかった」でした。「もういい」では止まらなかったのに、「それ、どうなったら終わりなの?」で止まった話。
- 続・「エコーチェンバー」 : 弁証法で描く信念硬直化のモデル構築記録
- 【Udemyコースレビュー】 Complete Generative AI Course With Langchain and Huggingface
- ⚙️自作アプリの管理をしやすくする❣️プロンプトビューアーと設定コンソール【Discord×Gemini 連載・第12回】
- 「次のモデルはヤバい」に、もう誰も驚かない。スレッドで一番伸びたのは、既視感だった
- お手軽LLMはじめてみた。その31 (Qwen3.8-27B、Qwen3-Coder-30B-A3B、Muse Glimmer-30B)
- LLMも人間も、次に来る言葉を見積もっている〜フィジカルAI⑧LLM〜
- DuckDBの開発元であるDuckLabsがAWS子会社になると発表。DuckDBはオープンソースのMITライセンスを維持
- 【対談連載 第五回】現場のドメイン知識を、数理モデルに翻訳する。アカデミアと事業を往復するデータサイエンティストが、リクルートで拓く「未耕地」
- セキュリティ初学者100名超が大集結!「しろおび夏祭り2026」で感じた学びと熱量