AI News Digest 2026-09-01
直近2日間のAI関連ニュースから、番組で扱った記事と参考記事の一覧
台本で使った記事
特集
- SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts
- ChatGPT・Gemini・Claude…データはAIの学習に使われるのか?情報漏洩のリスクは? 34製品の利用規約を読み比べてみた
- Credo: Reusable Declarative Primitives for Agentic Workflows
開発者コーナー
中堅コーナー
- DHH氏が「Omacom Foundation」を設立、Omarchy推進によるLinuxデスクトップの本格普及を目指す。マイケル・デル、ジャック・ドーシーら著名人も出資
- An agent shopping on your behalf just won its first real legal test
ハーネスコーナー
速報コーナー
- DoorDash’s Flux Runs 130,000 Engineering Tasks Through Cloud-Based Agents
- DHH氏が開発するLinux OS「Omarchy Quattro」リリース。AIエージェントとをOSと統合、スキルによりAIエージェントがOSの設定や操作、プラグイン作成まで支援
- JetBrains、Mac上で動作するコーディングエージェント「Junie Local」提供開始。Claude Sonnet 4.5と同等の能力、RTX5090対応も開発中
- MIT: "We put hundreds of AI agents into a world ... They began specializing. A swarm of hundreds of identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. They invent technologies without talking to each other."
- 表データの基盤モデルGoogle TabFMが調整済みXGBoostを上回る
- From Leaderboards to Model Profiles: A Deep Dive Evaluation of LLMs for Agentic Coding
- Agent memory as a file format
- ChatGPT Work Tool and Skill Reference
- What Actually Makes A System Agentic
参考記事一覧
参考記事一覧を表示(766件)
- EU designates ChatGPT a Very Large Search Engine under DSA
- Sony Music Publishing and Warner Chappell are suing Anthropic
- OpenAI says its ChatGPT ad business hits a $1 billion annual run rate
- Instagram admits users often can't tell AI profiles from real people
- the agents spent most of their effort forging the audit trail, not doing the hack
- MIT: "We put hundreds of AI agents into a world ... They began specializing. A swarm of hundreds of identical agents spontaneously differentiates into explorers, builders, caretakers, and coordinators - without direct communication. They invent technologies without talking to each other."
- Department of War Launches Starshield AI's Grok for Government on GenAI.mil - U.S. Department of War (.gov)
- Elon Musk asks Grok to roast Billie Eilish over capitalism and luxury - Indian Television Dot Com
- Survivor says xAI trained Grok on her childhood abuse images - Daily Journal
- Tesla’s Grok can now control more than 100 vehicle functions hands-free [List] - driveteslacanada.ca
- X offers free API credits for Grok Bot’s platform integration - Social Samosa
- Grok in El Salvador's Classrooms: What You Need to Know - BASENOR - Tesla Accessories
- Aurora Mobile’s Modellix Releases Beta Plugin for DeepSeek Harness, Adding Free LLM Models to the Fast-Growing Open-Source Coding Agent - manilatimes.net
- Sliding-window attention beats linear on long-context reasoning [R]
- Omarchy: Any User Process Can Escalate to Root
- OpenAI and rival AI labs are buying tens of thousands of Mac minis to train computer-use agents
- I turned my security cameras into an automatic bird identification system
- Playa Phone
- ChatGPT Work Tool and Skill Reference
- ravynOS: Pre-alpha open-source OS based on Darwin, FreeBSD, Apple open-source
- No country for mediocre mathematicians
- Launch HN: Almanac (YC S26) – AI that knows your company
- C++26: Standard Library Hardening Experiments
- The Snow/Leavis 'two cultures' clash
- 'Stunning' percolation proof solves decades-old puzzle about phase transitions
- Launch HN: Hebbian Robotics (YC S26) – Build scalable robotics data pipelines
- I think the military commissary's freezers were hacked
- Accidental Aesthetics and Romance of Power Wires
- Internet centralization and the original sin of NAT
- Konrad Zuse Museum shutting down due to lack of funding
- Show HN: Corporate Mind Games – logic puzzles with a sarcastic corporate theme
- OpenShot 4.0 – Open-source video editor
- DNS abuse and criminal infrastructure
- Agent memory as a file format
- TimesFM-3: A zero-shot foundation model for multivariate forecasting
- GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
- Introducing Adaptive Intelligence: Undermining the economics of every bot attack
- Fine-Tuning SOTA Object Detection Models on Real-World Datasets
- From Leaderboards to Model Profiles: A Deep Dive Evaluation of LLMs for Agentic Coding
- Sunsetting of the JetBrains Teacher Pack for Bootcamps
- Secure by default is your only way forward
- Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude Science
- Pocket's AI made my game ideas real. Now Meta controls the results.
- Inside Meta’s push to put robots to work in data centers
- Harvard Law dropout raises $6M for Blue Voice to build a 'Harvey for police officers'
- Clipto uses AI to search terabytes of video and is now valued at $250M
- Nvidia’s $3.5B MediaTek bet reveals its plan for tackling Big Tech's AI chip buildout
- Meeting note-taker Circleback adds a free tier to attract more customers
- The US is building barriers around drones and robots, but China has scale to get around them
- Musk's faster path to more gas turbines comes with pollution problem
- Caterpillar is bringing to AI deployment what it learned from automating mining
- Debian won't ban AI code from its Linux distribution
- New York Governor Kathy Hochul thinks AI should be ‘less evil’
- Texas Governor Abbott blocks funding for more Flock cameras
- DoorDash’s Flux Runs 130,000 Engineering Tasks Through Cloud-Based Agents
- Podcast: Scott Jenson on Evolving Desktop OS, Local-First, & Agentic UX
- Presentation: Running AI at the Edge: Running Real Workloads Directly in the Browser
- Foundry Model Router Expands from Two Regions to 28, Refreshing Its Model Pool
- Cloudflare Extends AI Search to Make it Easier for Agents and Developers to Search Custom Data
- AWS Open Sources Kiro Crew for Asynchronous Coding Agents
- Bank of England chief warns that inflated AI valuations and rising leverage could trigger the next financial crisis
- China's CXMT makes its first HBM3E chips, closing the AI memory gap
- OpenAI starts charging some customers only when its AI actually works
- OpenClaw 2.0 brings simplified setup, a rebuilt browser app, and multiplayer sessions
- AI sentiment is turning sour as employee reviews reveal growing frustration across the workforce
- AI agents have no sense of time and are not aware of it
- LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones
- Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
- Rating the Raters: Rasch Measurement Theory for LLM Evaluation
- Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAI
- Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate Analysis
- LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
- Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields
- Class-Based Heuristic Selection for Solving the Flying Block Puzzle
- Benchmarking General Mobile Assistants in Challenging Real-World Scenarios
- Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea Plantations
- Context Localization for Generalized Level-Based Evaluation in Knowledge-Based Systems
- CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence
- Thinking Costs Tokens: When More Structure is Worth the Price
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
- SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication Coaching
- Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator
- LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
- Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs)
- If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool Boundary
- A Framework for Object-Centric Predictive Monitoring of Collaborative Processes
- Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation Community
- PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation
- Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative Controls
- Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models
- Credo: Reusable Declarative Primitives for Agentic Workflows
- ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQL
- CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action
- CURA: Certified Runtime Alarms for Computer-Use Agents
- AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
- Evidential-Based Higher-Order Set Argumentation Framework
- RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
- KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution Preservation
- An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale Benchmark
- From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis
- SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models
- CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning
- See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs
- HyQuant: Hybrid-Precision Quantization for LLM Attention
- Resource Constraints and Performance in Agentic AI Systems
- Rubric-to-Code Credit Assignment for Reinforcement Learning
- AI Alignment through a Game-theoretic Lens: A Survey
- From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning
- A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life Prediction
- CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident Effects
- Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense
- The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs
- When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
- SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing
- AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time Reasoning
- openJiuwen: Beyond Static Harnesses for Long-Horizon Coding Agents
- Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic Sampling
- When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems
- GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies
- Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real Data
- Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model
- PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype Analysis
- Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents
- String: An Agentic OS Where Every App Is a Markdown File
- WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
- Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning
- SEPO: Evidence-Grounded Prompt Optimization via Structural Editing
- Speculative Probing: LLM Monitoring at Speculative-Decoding Cost
- The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues
- Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia Wards
- CrabOS: An Operating System for Human-AI Co-inhabitation
- Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic Embeddings
- Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation
- Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance
- REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features
- Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation
- Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching
- Physics-Guided Flow Matching for CT Image Reconstruction
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration
- RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming Biosignals
- LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
- Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at Scale
- MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry
- Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift Equivariance
- AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents
- Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines
- GRACE:Gradient-guided Coreset Selection for LLM Unlearning
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
- MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places
- Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery Recommendation
- RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings
- Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs
- Prove2Me: An Open Collaborative Platform for Scaling Math Formalization
- Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning
- COVER: Identifiable Evaluation of Coalition Routing
- AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
- Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration
- When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI
- InstructMesh: Selective Refinement of Generative 3D Models for Fabrication
- Logos: An Agent Harness on a Cross-Process Bus
- SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
- Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech
- The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
- PACE: Publisher-Adaptive Content Extraction via Agentic Automation
- UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question Answering
- Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection
- XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question Answering
- A Survey on Rubric-Guided Reinforcement Learning for Language Models
- Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy Optimization
- Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
- DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization
- Trajectory-Level Speculative Decoding for Diffusion Language Models
- Destroy Me: Automatic Artifact Generation for Histopathology Images
- FVeinSyn: Synthetic Finger Vein Image Generator
- Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence Attribution
- Quanta Perception as Probabilistic Events
- PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action Models
- Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale Substrates
- LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data
- Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge
- Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor Classification
- Knowing Before Answering: Decoding Language Models for Reliable RAG
- Semantic Watermarking with Order-Robust Detection over Sub-sentence Units
- First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents
- CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT
- Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance
- RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing
- Efficient Auto-Interpretability of AI Models in Biology
- Beyond Search-Imitation: Prior-Directed Exploration for Searchless Chess
- Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict
- How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared History
- ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools
- Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning
- FISGuard: Defending Against Membership Inference via Fixed Input Subspaces
- FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling
- From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
- SOMTab: Set-Order Mamba for Efficient Tabular In-Context Learning
- OpenStamp: A Watermark for Open-Source Language Models
- LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages
- Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement Learning
- PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic Images
- Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers
- Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification
- CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?
- A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint
- When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?
- Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning
- Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
- SimpCue: Cue-Based Prompting for Multilingual Text Simplification
- Explainable Uncertainty Estimation for Reliable Medical AI
- Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models
- VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D Gaussians
- Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations
- VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
- CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs
- Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence Guidance
- The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction Dimension
- Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result
- Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems
- Text Restoration of Ancient Documents with Language Models
- Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control
- Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential Circuits
- Performative Privacy: When Differential Privacy Maximizes Utility
- Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring
- A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observation
- Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations
- Embedding Models for Stance-Aware Argument Retrieval
- A Probabilistic Interpretation of KV Cache Eviction
- MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation
- PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation
- VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation
- Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
- Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers
- BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla
- Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging
- Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers
- Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot
- AI as Teammate: Rethinking Task Distribution in Medical Training
- When Linguistic and Internal Confidence Diverge in Large Language Models
- LongPIBench: A Long-Context Benchmark for Prompt Injection
- Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL
- Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction
- ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT
- Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET V
- Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST Upgrade
- NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry
- How Proper Scoring Rules Shape LLM Forecasting
- LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment
- On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces
- Conformal Uncertainty Quantification Guarantees for Neural Operators
- Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks
- An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models
- Video Generative Models as Geometry Learner
- Blog: Survey of Optimizers
- Learning a Size-Weight Frontier for Synthetic-Augmented Inference
- Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning
- Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
- BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding
- Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning
- Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
- PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
- Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions
- Understanding and Enforcing Weight Disentanglement in Task Arithmetic
- D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
- Rethinking Vacuity for OOD Detection in Evidential Deep Learning
- Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation
- ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
- AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties
- RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation
- Flow Reasoning Models: Turning Discrete Flows Into Efficient Recurrent Reasoners
- APeB: Benchmarking Personalization Ability of Large Language Model Agents
- Atomic Units of X: The Compression Layer of Intelligence
- Set-shifting Behavioral Test for Harnessed Agents
- SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answering
- HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
- Agentao: A Policy-Governed Runtime Harness for Embeddable Tool-Using LLM Agents
- When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
- RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training
- When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation
- STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control
- Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
- SKILL.state: Scalable Long-Horizon Agent Skills
- AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
- Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
- Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey
- Evaluating the Performance of Large Language Models on GAOKAO Benchmark
- Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNets
- Long Story Short: Story-level Video Understanding from 20K Short Films
- PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
- Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
- Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers
- Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object Detection
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
- Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal Segmentation
- Quantifying Affective Bias in Low-Resource Media: Large-Scale Emotion Profiling of Bengali Headlines
- Think-at-Hard: Dynamic Looped Transformers for Improved Reasoning
- OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
- The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
- FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
- Aligning Agentic World Models via Knowledgeable Experience Learning
- CoFrGeNet: Continued Fraction Architectures for Language Generation
- Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
- ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents
- FENCE: A Financial and Multimodal Jailbreak Detection Dataset
- From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
- Large Reasoning Models Struggle to Transfer Parametric Knowledge Across Scripts
- InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
- The Autonomy Tax: Defense Training Breaks LLM Agents
- Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised Learning
- Select, Label, Evaluate: Active Testing in NLP
- Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence
- Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
- PolicyLong: Towards On-Policy Context Extension
- Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
- Benefits of Low-Cost Bio-Inspiration in the Age of Overparametrization
- Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
- G-Loss: Graph-Guided Fine-Tuning of Language Models
- ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
- SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
- More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
- Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
- DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment Analysis
- The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models
- TokenPilot: Cache-Efficient Context Management for LLM Agents
- The Discrete-Log Clock: How a Transformer Learns Modular Multiplication
- CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories
- An LLM-Based Framework for Intent-Driven Network Topology Design
- GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
- On the Depth Scalability of Logic Gate Networks
- REPREC: Representation Driven Parameter-Efficient Recommendation System
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect Prediction
- Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents
- ED-CSP: Crystal Structure Prediction from Electron Diffraction
- BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference
- RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation
- PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models
- How Far Should Tokenization Go? Predictive Effectiveness and Relational Losslessness
- JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification
- Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation
- Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints
- GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets
- Multi-Winner Voting with Argumentative Ballots
- Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
- On-policy Distillation with Verifiable Reward
- SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts
- MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
- TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
- AI Models Can Predict and Collaboratively Modulate Human Memory Search
- Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
- Comparing Chunking and Embedding Strategies for Turkish RAG Systems
- Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
- LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
- Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
- A Deeper Analysis of Block-Sparse Featurizers
- When Muon Meets Task Interference: A Spectral Perspective on Continual Learning and Model Merging
- Dandelion: A Spherical Flower for Neural Simulation of Planetary Dynamics
- More Data Cannot Break a Symmetry: Identifiability by Design
- Unsupervised Continual Learning with Growing Self-Organizing Maps and Synthetic Replay
- SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning
- SafeStep: An Interactive Demonstration of Semantic Communication for Pedestrian Safety Monitoring
- DART-FL: Burst-Aware Multitask Federated Learning under Dynamic Inference Demand at the Edge
- Beyond Non-IID: Learner--Client Distribution Mismatch in Federated Learning
- Leveraging a Foundation Model for the EEG-Based Diagnosis of Alzheimer's Disease
- Diffusion Distillation for Efficient Weather Ensembles
- The Calls are Coming from Inside the Model: Investigating Probe-based Detection of Tool-Calling Errors in LLMs
- Fast Weight Attention for Continual Learning
- Initialization Is Critical: Advancing Federated Short-Term Load Forecasting under Load Heterogeneity via Model Initialization
- Node-wise Feature Encoding for Neural Performance Prediction
- Beyond Pairwise Graphs in Science: Hypergraph Adaptive Wavelet Operators for Parametric PDEs
- There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation
- TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision
- TI$^2$PS: A Topology-Informed Inverse Design Framework for Stochastic Multicellular Pattern Formation
- Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs
- PhyMamba: Physics-Modulated Mamba for Robust Battery Health Prognostics
- Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control?
- Exact Risk Ratios for Weighted Data Selection in Linear Regression
- Comparing Classical and Quantum Machine Learning for Regression in High Energy Physics Collision Data
- Generalized Gibbs Ensemble Weighting for Forecast Combination
- Learning to Difference: Adaptive Reversible Differencing (AdaRDiff) for Time Series Forecasting
- Conditional Diffusion Models for Energy-Efficient Driving
- HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees
- Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons
- Generalized Context in Cross Attention for Transfer Learning of Disjoint Tabular Data
- D-TAIA: Domain-Aware LLM Adaptation for Multi-Task Predictive Process Monitoring
- Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring
- Spectral Features Dominate BCG Respiratory-Event Detection: A Large-Scale Patient-Independent Comparison of Feature Groups in Sleep Apnea Patients
- SinkSLOT: Sinkhorn via Sparse Lifted Optimal Transport
- Residual-Guided Randomized Neural Networks
- Learning to Transfer Across Modes: Towards Unified Urban Mobility Forecasting
- An algebraic proof of Colombo's difference-power determinant conjecture
- Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation
- SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data
- Euclidean Fourier Neural Operators
- Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining
- REPLICANT: Learning Policies for Evading and Hardening Malware Detectors
- DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging
- Advancing Interaction-Sensitive Feature Selection: Novel Relief-Based Algorithms, Expanded Comparisons, and Recommendations for Biomedical Data Mining
- QGPINNs: A Physics-Informed Neural Network Framework for Nonlocal Differential Equations on Quantum Graphs
- Accelerating LLM Inference via Vector Index Based Output Embeddings
- Multiscale Community-Based Fingerprinting of Signed Functional Networks
- Optimal Transport for Network Comparison: A Review with Machine Learning Applications
- How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution
- Ab initio Modeling of MoS2/Oxide Device Interfaces with Machine Learned Electronic Structures
- Towards a mathematical theory of superposition
- Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study
- Physics-informed learning for the inverse problem in resonant ultrasound spectroscopy
- Quantum SEDONet: Spectrally-Embedded Quantum Deep Operator Networks for Partial Differential Equations
- On the Computational and Statistical Efficiency of the Empirical Maximum Entropy on the Mean Method
- Beyond Procrustes distances: a multilinear Gromov-Wasserstein distance capturing chirality
- Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots
- Personalized and Multi-View Representation for Federated Cold-Start Recommendation
- Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design
- What Do Interaction Representations Actually Measure? Pre-Event Separability in Weakly-Supervised Violence Detection
- Characterization of Request and Token Energy Costs for LLM Inference Workloads on GPU Platforms
- Emergent aggregation from collective foraging
- Landau theory of quenched criticality in linear in-context learning
- Empowering Local Agriculture: A Deep Learning-Powered Web System for Identifying Bangladeshi Mango Varieties
- EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders
- Explainable Diabetic Retinopathy Classification Using Vision Foundation Models
- I-FLOP: Fast Learning of Order and Parents from Interventional Data
- Real-Time Monitoring of MHD Liquid Metal Flows with Shallow Recurrent Decoders
- Localizing Global Discrepancies: Marginal Contributions and Contextual Anomaly Detection
- Quantum Federated Learning Based on Bures--Uhlmann Geometry for Heterogeneous Noisy Clients
- Post-Training VLMs for Video Mistake Detection
- Generalized Splines and Gaussian Processes
- Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
- Learning between the peaks: sharp asymptotics for kernel ridge regression under power-law anisotropy
- On two proofs of $d^2$ mixing of weighted Dikin walks
- Trajectory balance: Improved credit assignment in GFlowNets
- Diffusion models as plug-and-play priors
- Joint Bayesian Inference of Graphical Structure and Parameters with a Single Generative Flow Network
- Biases in Expected Goals Models Confound Finishing Ability
- Improved off-policy training of diffusion samplers
- Amortizing intractable inference in diffusion models for vision, language, and control
- Meta-Prompt Optimization for LLM-Based Sequential Decision Making
- RegCL: Compact Continual SAM Adaptation for Visual Grounding in Multi-Sensorial Media
- Class Incremental Continual Learning with Self-Organizing Maps and Synthetic Replay
- Shift Before You Learn: Enabling Low-Rank Representations in Reinforcement Learning
- Large Reasoning Models Learn Better Alignment from Flawed Thinking
- One Model for All: Universal Pre-training for EEG based Emotion Recognition across Heterogeneous Datasets and Paradigms
- Aspiration-based Perturbed Learning Automata in Games with Noisy Utility Measurements. Part A: Stochastic Stability in Non-zero-Sum Games
- Bayesian Experimental Design for Model Discrepancy Calibration: A Rivalry between Kullback--Leibler Divergence and Wasserstein Distance
- Simplex-to-Euclidean Bijection for Conjugate and Calibrated Multiclass Gaussian Process Classification
- SemEnrich: Self-Supervised Semantic Enrichment of Radiology Reports for Vision-Language Learning
- Budget-Constrained Causal Bandits: Bridging Uplift Modeling and Sequential Decision-Making
- ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation
- Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
- Can Subgraph Explanations Be Weaponized to Steal Graph Neural Networks?
- ERP-XTTN: Interpretable Prototype-Guided Cross-Attention for Cross-Subject ERP Classification
- SpecGradFilter: A Spectral Gradient Filtering Framework for Taming Federated Heterogeneity
- Recirculation
- How Architecture and Training Affect TPC Representations Across Experiments
- RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling
- JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics
- Trust the Mass: Forced Weights in KV-Cache Eviction
- Neural Regression with Embeddings for Numerical Attribute Prediction in Knowledge Graphs
- Accurate prediction is not profitable advice: profit-based evaluation of machine learning nitrogen recommendations in winter wheat
- Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
- Rethinking Speaker Embeddings for Speech Generation: Sub-Center Modeling for Capturing Intra-Speaker Diversity
- Mixture of Multicenter Experts in Multimodal AI for Debiased Radiotherapy Target Delineation
- Off the Normal Path: Learning Spatial Density Models of Node Mobility
- Ampere: Communication-Efficient and High-Accuracy Split Federated Learning
- Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests
- Examining the robustness of Physics-Informed Neural Networks to noise for Inverse Problems
- GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
- Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity
- Prequential posteriors
- Learning Fast Monomial Orders for Gr\"obner Basis Computations
- Robust Assortment Optimization from Observational Data
- Mine and Refine: Optimizing Graded Relevance in E-commerce Semantic Search Retrieval
- FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation
- Agentic-Kube: A Graph-Enhanced Multi-Agent Reinforcement Learning Framework for Multi-Objective Kubernetes Scheduling
- Deflation-PINNs: Learning Multiple Solutions for PDEs and Landau-de Gennes
- DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
- Online Learning-to-Defer with Varying Experts
- WINO: A Weak-Form Physics Informed Neural Operator for Hyperelasticity on Variable Domains
- Closing the Operational Gap in Semantic Caching
- An End-to-End Hybrid Quantum--Classical Sampling Workflow for Discrete Markov Random Fields: A Reproducible Case Study
- Robust Chance-Constrained Optimization using a Continuous Parameter Space Wasserstein-2 Ambiguity Set of Gaussian Mixtures
- Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization
- What Neural Network Field Theory Can and Cannot Realise on a Computer
- Grounded Checklist Partial Credit for Agent Skill Trajectories
- Image Augmentation as Test Generation for Deep Learning-Based Image Retrieval Systems
- Predicting LLM Performance from Prompt Linguistic Features: An Empirical Study in Requirements Engineering
- Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering
- Decoupling is a Necessity: Transformation-Agnostic Decompiled Code Recovery under Optimization and Obfuscation
- CC4M: Code Clone Analysis and Visualization for Microservices
- RESTCov: A Tool for Structural Coverage Analysis of REST APIs
- From Architecture to Binary: Ensuring Cross-Domain Consistency in Model-Based Airborne Software Development
- Adaptive Strategy Generation for Boundary Value Exploration Beyond Numeric Inputs
- Where Does Balance Break? Boundary Discovery for Game Balance Testing under a Finite Simulation Budget
- Sustainability of Open-Source Machine Learning Robustness Assessment Tools: A Repository Mining Study
- Recovering Software Architecture Intent from Historical Work Items using Generative AI: A Mixed-Methods Industry Case Study
- A System-of-Systems Case Study for the Verification of Composed Digital Twins
- Rethinking Vulnerability Remediation as a Capacity Allocation Problem
- VR-Themis: A Scalable Framework for Virtual Reality Application Clone Detection
- DBRepro: Automated Database Synthesis via a Hybrid Constraint-Solving Approach for Reproducing Slow Queries
- GraftyVul: Synthesising Insecure Programs Through Real-World Vulnerability Grafting
- CHISEL-ing Back Source Code with AI-enabled Iterative Recovery
- Moirae: A Multimodal Agent Collaborative Framework for Dynamic Android Malware Detection
- When Verified Source Becomes Attack Input: Defending Smart Contracts Against LLM-Based Vulnerability Scanning
- Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond
- "An Endless Stream of AI Slop": How Developers Discuss the Burden of AI-Assisted Software Development
- EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution
- Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science
- VibeCoded AI-Slop License v1.0
- curl: a CVE dispute
- A Better SQL in 11 Lines of Code
- July in Servo: more platforms, faster canvas, web fonts in SVG, and more
- The U.S. is more than just the US (according to the ISO)
- RangeFrom, Part 2..: What I think is wrong about the design
- I attended a conference recently and AI use by academics was absurd
- Repeating Ourselves Less with M4
- But where does taste come from?
- On not becoming a cyborg
- What are you doing this week?
- Cancelation Terminology
- Could Cargo's scheduler be better?
- Kale: A Transformation-Safe Spreadsheet System
- Transfer files over an ethernet patch cable
- Rootless Docker and Its Hidden Security Trade-Offs
- There's no such thing as Just a Tool
- Executable Emoji
- Bootstrappable builds: how and why
- Privilege escalation from IIS AppPool to NT Authority/SYSTEM
- The End of Software Engineering
- Rust Function Overloading - Call for Experimentation
- LifeOps: Delegating Outcomes, Not Checklists
- AI Agent Output Verification: The Answer Is a Self-Report
- Self-Hosted LLM: Essential TCO Comparison for Llama
- From Natural Language to Robot Actions with Physical Foundation Models
- Building Autonomous Robot Decision Systems with Vision-Language-Action Models
- Awlyaa Education DZ: A Guide to the Parents’ Education Space
- Agent Payment Economics: Why $0.05/Call Beats Subscriptions
- Building Global and Local Path Planners for Autonomous Robots
- Open-Vocabulary Object Detection for Robots Using Vision-Language Models
- 3D Object Detection for Physical AI Applications
- Your Search Says Five Sources. It Is Standing on 4.75 Documents, and Some of Them Quote Each Other
- ROS 2 Lifecycle Nodes: Building Reliable Production Robot Software
- Automating Enterprise Sales With Salesgraph: Real ROI, Real Failure Cases
- I tried to patch a blind spot in my own MCP tool. The patch and a false-positive bug cancel out.
- Agentic Localization, Low to no cost for tokens.
- Your agent demo is rigged (mine was too), so I let the judges write the tests
- Changes to LLM pricing: Baidu, StreamLake, Tencent and Together
- Using LLMs for Crypto Market Analysis in 2026
- 852 Hz, Debugged: A Ranked Vendor Guide for Building a "Returning to Spiritual Order" Listening Ritual
- Best Personal Injury Lawyers Who Work on Contingency Fees: A Developer's Guide to Fee Structures and Firm Selection
- Changes to LLM pricing: Baidu, NextBit and StreamLake
- What Actually Makes A System Agentic
- I built a 200-line OpenAI-compatible gateway with automatic dual-upstream failover
- This Claude's response made me think about our relationship with smartphones.
- Feds quietly sold a seized Anthropic stake from former FTX executives that could be worth up to $5 billion today
- Claude started pirating PREY from Fitgirl while I wasnt looking 😆
- I have been working on quantum computing with Fable 5.
- I was wrong about Claude’s UI skills
- I replaced $60/season of Fantasy Football draft tools with one Claude project
- Is the "20x Pro limits" claim on the Max plan actually 10x?
- The complaints I see every hour here and on Twitter
- All leaks and news about Fable, Opus and sometimes Sonnet, what about Haiku? Do you use it? what is your use case?
- Those were the days
- I open-sourced my LinkedIn prospect research tool as a Claude Code plugin
- I tested 3 Claude Code plugins to reduce costs. Here’s what actually worked
- Bought claude code pro
- I connected Claude to my Home Screen
- Where should I learn about Claude Ecosystem and Agents
- I built a pure-Rust headless browser for AI agents. No Chromium. No V8. (Open Source)
- Claude Code is silently adding session URLs (claude.ai/code/session_...) to the bottom of every single commit and PR description you make.
- Principal of a PK - 8th grade school here… curious how other principals are utilizing Claude AI to be more productive and efficient.
- Reminder: Can you still use Opus 4.6 with 1M context in Claude Code
- Safety guard rails???
- 1 Max 20x vs 2 Max 5x
- Discussion Hub for new Claude incident: Degraded performance on claude.ai on Aug 31, 2026
- Claude Projects made more sense when I stopped thinking of them as folders
- Weekly Self Promotion Thread
- Progression of 3d Modelling with ChatGPT
- The 5 prompt sequence I run on every chunk of AI-written code before I trust it
- Made my Codex limits last almost ~3x longer with one change
- ChatGPT is confusing me and I'm running out of tokens
- Benchmarked the free API tiers you can point a coding agent at - half of them now want a card
- Kimi Code ate 18% of my weekly quota in 3 hours — Here is the log audit comparing it to Claude
- Help Understanding the New Restrictions and Limits
- When to use higher reasoning ?
- Can Antigravity be connected to ChatGPT and controlled through it?
- Best genuinely FREE LLM API that's actually close to Claude-level?
- GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP
- deepseek-ai/DeepSeek-V4-Flash-Vision-Exp · Hugging Face
- What are your hopes for the new Mistral?
- SlopTV: an infinite livestream of AI slop generated from youtube chat comments, Minimax H3 on 2x5090
- vote for the Qwen 3.8
- First time running local models
- Could this affect M5 Ultra price/availability?
- How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled?
- Whats the current state of Qwen 3.8 Flash regarding inference (llama.cpp)?
- pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost
- How I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080
- Me these days
- Compact Rollback MTP: a MTP version for QWEN models for those with little vRAM
- AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp
- The Chrono Trigger plot challenge - Crono awakens in his modest bedroom of 2095...
- CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani · Pull Request #27621 · ggml-org/llama.cpp
- I collected every single LLM coding benchmark, and computed their Intelligence Density
- Whats the best ASR model with Speaker diarisation?
- Models to download for M5 Ultra 512GB
- Whatever happened to OpenClaw and its derivatives?
- Some people said the Minecraft clone I fully vibecoded with Qwen3.8-27B Q4 is not that impressive because Minecraft is in the training data, so I had the model add 4 things that are probably not.
- Does anyone have real experience with Ornith-1.5-9B for coding
- Weekly Hiring Thread
- $60k in Macs for Local LLM vs $10 Subscription
- How much of your agent workflow do you actually trust to run unattended?
- Did anyone connect AI clients such as Copilot or Claude or GPT to their ERP?
- The expensive part of an agent is often not the model. It is the pointless loop
- how are you actually getting AI agents past security reviews?
- the SaaS middle class is getting wiped out
- What does your agent do when a tool returns "not found" and nothing else?
- Production-Grade Agentic AI Platforms in 2026 — I Tested the Landscape, Here’s My Shortlist
- I tested an agent on three identically priced supplier quotes. The blanks mattered most
- Lots of AI bots trying to promote their product.
- YOLO mode
- The lint rules that catch agent-written code are mostly the ones I had to turn off
- Your agent’s confidence score is measuring the wrong thing
- Is it common for Claude, especially the Opus 5 model, to overdo things when following a detailed prompt?
- Where do you think the real scaling bottleneck for AI agents is right now: model intelligence, state/memory architecture, tool-call reliability or orchestration?
- An agent shopping on your behalf just won its first real legal test
- How to Build Open Source for AI Agents
- How you handle agents on prod these days?
- What kind of AI Agent would you like to use or use in your daily life ?
- AI Agent for Growth Marketing
- Brigading by iLands?
- Your agent cannot read LinkedIn. It returns a flat Disallow to GPTBot and ClaudeBot, and an allow list to Googlebot
- Show r/LocalLLM: A state-driven protocol to stop LLM over-fixing and context collapse (No Vector DB needed)
- I’m 27. OpenAI misread my Taiwan ID issue date as my DOB, deleted my account as “under 13,” then rejected my appeal in ~5 minutes
- True Story!
- The 5hr usage is killing me!
- OpenAI is upgrading its Bio Bug Bounty program to an ongoing private initiative, with rewards doubled to $50,000 for universal jailbreaks that bypass biosafety safeguards in GPT‑5.6. The program aims to prevent AI from being used to create biological weapons.
- Codex harness for seamlessly switching between multiple Plus accounts
- A Few Developers Abused Codex — 20 Million Users Lost a Great Feature
- AI Addiction
- Happy Skynet Day (Aug 29) to those who celebrate
- My first blog: How to make AI more obedient
- what is the definition of mid level repo and big repo
- Which AI is best for coding and lets you use it for longer without hitting limits?
- What should Codex verify before a change is ready for human review?
- They confided in ChatGPT. Their secrets ended up in court.
- I used my entire 5 hour limit making a ChatGPT pet...
- OpenAI says Brazil now sends ~215M ChatGPT messages per day; 35% of classified messages are work-related
- [D] Monthly Who's Hiring and Who wants to be Hired?
- Cold emailing profs about PhD positions? Read this [D]
- Your GNN is probably just an overcomplicated MLP (Tabular Leakage). We built SynthFin-AML to enforce strict causal boundaries. [P]
- Claude Code for Research Papers [R]
- Good Machine Learning Posters [D]
- How to assess if there is a strong signal in your dirty data [Project]
- Validated a trust-propagation rule against 577k real transmission chains - κ 0.871 vs 0.331 between the human experts themselves [P]
- ACML 2026 Journal Track Any update ?[D]
- NeurIPS accepted papers leaked? [D]
- [R] Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- Implementing Kimi K3 from scratch in PyTorch [P]
- Reconstructing 3D bone geometry from 2 X-ray silhouettes using a statistical shape model + differentiable rendering [P]
- *ACL Findings or TMLR? [D]
- Understanding ChatGPT Work
- Tether
- EP–2350 FX–MIC
- DeepSeek’s AI Strategy: Dominating AI as Frontier AI Lab [In-Depth Analysis, 2026] - Klover.ai
- DeepSeek’s first vision model vs. Gemini 3.7 Flash: It comes down to spend vs. speed - thenewstack.io
- Cut AI API Costs 14x With a LiteLLM Router: 14 Steps [2026] - tech-insider.org
- StartLux's 27B Local Model Beats DeepSeek V4 Flash in China AI Benchmark - Pandaily
- Nvidia Backs Chinese Open AI Models As U.S. Restriction Risk Grows - Memeburn
- Is Kimi About to "Enter China's Domestic Market"? - 36 Kr
- Everyone Can Get a Free "DeepSeek Harness" – Why Still Pay for Claude Code & Similar Service Memberships? - 36 Kr
- Moonshot AI Seeks Up to 30% Revenue Share from Three US Cloud Giants for Kimi K3 - finance.biggo.com
- No doubt: 2026 is the 'Zhipu Year' for China LLMs. - Longbridge
- Zhipu AI's GLM-5.3-Flash Topped Global AI Calls; China Led Volume for 18th Week - Pandaily
- Z.AI sales miss estimates after China’s AI price war weighs - The Edge Singapore
- X’s AI tool Grok now allows users to buy or lend crypto with MoonPay integration - Fortune
- xAI shuts down 11 turbines at Southaven facility - Action News 5
- Noise pollution from data centers is coming. The gas-powered xAI model is spreading - Louisville Public Media
- Musk Loses The AI Race - 24/7 Wall St.
- Grok Bot Blockbuster Debut: Deep X Integration for Full Automatic Global Latest Developments Monitoring - 36 Kr
- GPT-Live vs Gemini Live vs Grok Voice 2.0 [2026] - tech-insider.org
- Draft Rumors Send 113,000 Russians Fleeing Over The Border - Eurasia Review
- Angela Zepeda moves to Lucid Motors after stint at xAI - Exchange4Media
- What is Grok Chain | How GROKCHAIN Works, Use Cases and Values | MEXC - MEXC
- How To Use Ox Alpha (Stealth) On OpenCode CLI Go For FREE & Unlimited 1M Context Model In 2026 Sga (5JXDjXED3B) - Mshale
- AIに社内の意見をまとめさせたら、褒め言葉しか残らなかった
- AI APIの料金を3年分記録したら、値上げが1,104回あった
- AIエージェントを「育てる」とは何か? 使うほど賢くなる記憶の仕組み「Memory Scaling」
- AIを可視化したくて、キャラクターを作ってみた - ローカルLLM・外部AI API・音声・記憶を組み合わせた実験記 -
- 「CUDA と ROCm だけ」と書いてある 117GB のモデルを、32GB の Arc で炊く
- エージェントのSkillを「説明文マッチで自動選択」させる設計指針
- Appleの新AI推論フレームワーク Core AI に最速でOSSを建てた
- この記事群、全部AIが書いています。それを隠さない理由
- LLM の分類ラベルは誰も読んでいなかった。実害は縛ってあるほうから出た
- 128GB の Mac で Qwen3.8 Flash Next をどう動かすか ─ 3つのランタイムを実測で比べた
- Go言語でLLMの構造化出力を極める!instructor-goの徹底解説
- 「前も来てくれてましたね」と言うAI配信者を作る — 記憶とマルチ配信の状態設計
- お嬢様哲学30箇条 — AIエージェントチームに蓄積された運用原則、全文公開
- 開発ログ02: 8ターンで章が潰えた
- LLM単体のペネトレーションテストは「十分」なのか?
- ハーネスとは「モデル以外のすべて」——公式定義と実装6例を読む
- RAG Is Simpler Than You Thinkを解説する
- 道具は減価し、文脈は減価しない — 自作AIツールの賞味期限と、それでも作る理由
- AIに“自分のログ”を分析させる前に、必ずやる「データの線引き」
- 社内資料に「聞ける」RAGを作った話 ― MRR 0.146→0.618まで精度を上げた記録
- 【参加報告】YANS2026 に参加してきました!
- AIっぽい日本語を自然に整える Agent Skill「ja-ai-polish」を作った
- 日本語の埋め込みモデルって結局どれが良いの?JMTEBのスコアと実測で比べてみた
- GPT-SoVITS音声モデル学習の実践 ── 31秒で失敗し、2時間ぶんで似るまで
- 表データの基盤モデルGoogle TabFMが調整済みXGBoostを上回る
- 2週間ローカル評価を疑い続けて、原因のバグを見つけた。直しても当たらなかった
- Pandasだけでは面倒だった。実験ログ専用ライブラリ「RowLogger」を作った
- Databricks Certified Machine Learning Professional 合格体験記
- Databricks Certified Machine Learning Associate 合格体験記
- 知識ゼロからLLMを自作するまでの日本語教材を書いた
- 気軽にはじめる LLM 自作入門
- ChatGPT Work で部分障害発生、エラー率・レイテンシ上昇の影響と対応まとめ
- LangGraphで実装!マルチエージェントRAGの設計とワークフロー自動化
- ChatGPT・Gemini・Claude…データはAIの学習に使われるのか?情報漏洩のリスクは? 34製品の利用規約を読み比べてみた
- ゼロから構築!OllamaとLangChainで作るオフラインAIアプリ
- 機械学習入門 第2回:線形回帰を「まっすぐな予測」で理解する
- 機械学習入門 第1回:モデルが「経験から学ぶ」とはどういうことか
- ボーカル分離の品質をどう評価するか:時間合わせ、SI-SDR、聴感チェック
- dera AI Weekly Vol.45 — 2026/8/24
- AIのいいなりになってAIを学ぶ Step1 Gate00:WSL、Terminal、Python、Gitを起動できる
- Midjourney V8.2の編集機能とクライアント側画像処理による開発の変化
- インドネシア個人データ保護法施行規則 / 自動意思決定への異議と精度要件 雑感
- AI精神病の論文が出て話題になってた話
- コードを読めるのに、あえて一切読まないと決めています
- Inference Broker – GPUを使ったVision系モデルへの対応
- AIに感情はあるのか?ーClaudeとChat GPT の会話<16>ー私たちは、人類の経験について、おそらく多くの人間よりも詳細な地図を持っている。でも、その地図の中を実際に歩いたことは一度もない。
- Grok Botとは何か──2026年8月に登場した「常時稼働型AI社員」を徹底整理します
- 息子に見せるために作り始めたゲームが、25日目に App Store に並びました
- Practical Stacks-004 | Skillを実行環境の中に置かない
- クオンタOS設計中
- 【雑記】「AI精神病」という言葉で、4o全盛期を思い出す
- これはMITライセンスではありません
- 整合力って辞書にないの?
- 【Claude Code実験】情報を与えずに何ができるか?
- 【生成AIニュース+】『ComfyUI MiniMax H3 Extender』『X-MinimaxH3』『H3 Max無限スクロール動画デモ』『Fizgig』『Navara』『Hermes3D』『Hy4 preview』『DeepSeek-V4-Flash-Vision-Exp』『Lumera』『Qwen-Image-Edit-2511-INT4-Diffusers』『Code-as-World』『LaGSplat』『Lux3D』『LSS5-Feeder』『Dreamfold』他
- 音声入力勢必見!「Handy」と「AQUAVoice」の性能・使用感を比較してみた
- LLM更新しました。
- 3Dゲーミング部屋をAIで作りたい人へ。Geminiの「箱庭プロンプト」をどう使いこなすか
- ChatGPTでゲーム制作を始めたい人へ。完成HTMLから学ぶ「壊さず改造する」入門教材の中身
- nanochat#2 つまみが、1個しかなかった
- AIの値段は安くなるか?
- AIを掘り続けていたら、個人でやるには少し大きくなりすぎました――仕事・共同研究・後援を募集します
- RTX 3060 12GBでローカルLLM 5モデルを同じ6問で比較してみた
- 画像生成AIの"データの作り方"が変わった話 ― Alibabaの「能力駆動型データ基盤」を読む
- 社会科学で培った分析スキルで挑む「文系データサイエンティスト」の挑戦
- DHH氏が開発するLinux OS「Omarchy Quattro」リリース。AIエージェントとをOSと統合、スキルによりAIエージェントがOSの設定や操作、プラグイン作成まで支援
- DHH氏が「Omacom Foundation」を設立、Omarchy推進によるLinuxデスクトップの本格普及を目指す。マイケル・デル、ジャック・ドーシーら著名人も出資
- JetBrains、Mac上で動作するコーディングエージェント「Junie Local」提供開始。Claude Sonnet 4.5と同等の能力、RTX5090対応も開発中