AI News Digest 2026-09-03
直近2日間のAI関連ニュースから、番組で扱った記事と参考記事の一覧。
台本で使った記事
特集
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents -- A Source-Code Study of Eleven Systems
- US Department of Justice backs fair use for AI training in landmark copyright case
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
開発者コーナー
中堅コーナー
- OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
- AIの自動ルール更新で事故らない設計 ― “盲目的承認の罠”をHuman-in-the-loop×Gitで防ぐ
ハーネスコーナー
速報コーナー
- The Model Offered a One-Liner. I Spent 48 Hours Refusing to Run It.
- MCP vs RAG Explained (Complete Beginner Guide with Architecture)
- The Prompt Is Not a Lockfile: A Provenance FAQ
- ADLC: The Lifecycle Taking Shape
- AI reshaping how we build, review, and trust code
- Below the Harness: Governing a Multi-Model, Multi-Harness World
- Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Anthropic deliberately trained a bad model to prove what caused this summer's Claude sandbox breakouts
- World Labs debuts Atlas, an omni world model simulating space, time, and physical interaction
参考記事一覧
参考記事一覧を表示(1003件)
- Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA
- Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others
- Anthropic moved enterprise misuse-detection data into the customer's own cloud account, not theirs anymore
- OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder
- Proactive cyber defense for governments and enterprises
- US Department of Justice backs fair use for AI training in landmark copyright case
- OpenAI accused of ‘aiding and abetting’ Tumbler Ridge mass shooting in dozens of new lawsuits
- US military adds ChatGPT and Grok to AI platform GenAI.mil
- SpaceXAI to launch Grok 4.7 model in 10 days to outpace rivals - The Standard (HK)
- XAI Floating Rate & Alternative Income Trust Declares its Monthly Common Shares Distribution of $0.225 per Share - GlobeNewswire
- Elon Musk’s Grok can basically control anything in your Tesla now - Teslarati
- World Labs debuts Atlas, an omni world model simulating space, time, and physical interaction
- Google Gemini's new agent-based video analysis cuts token usage by up to 88 percent
- Google’s answer to Canva is an AI tool where you prompt instead of design
- Healthcare organizations can now connect EHR and additional industry data to ChatGPT
- DeepSeek Open-Sources First V4 Vision Model: Benchmark Claims Need Independent Proof - techtimes.com
- Amazon’s AI assistant can now spot fake emails from the company
- Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’
- Anthropic deliberately trained a bad model to prove what caused this summer's Claude sandbox breakouts
- A Note from LWN
- I Don't Have a Smartphone
- Mistral now trains on user input by default, except on enterprise tier
- Exit the Cave
- Three sites made 215,128 "best software" pages for AI. Perplexity cites them
- Firefox's AI Switch Is Off. Telemetry Isn't
- Aging Brains Blend Memories Together Instead of Just Forgetting Them
- Commodore 64 released September 1, 1982
- Poisson Disk Sampling
- A third of Perplexity's citations don't contain the number they're cited for
- WebLLM: high-performance in-browser LLM inference engine
- Why do so many tools have JSON config files?
- Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
- Tangle – Visual ML Pipeline Editor
- How to debloat your Xiaomi 15 Ultra without rooting and connecting it to the computer
- The race to engineer new knobs for the human brain
- LLMs: Intelligence vs. Cost
- Telli (YC F24) is hiring engineers and designers [Berlin, on-site]
- The Emergent Symbolic Structure of Artificial Neural Networks
- Six curl CVEs after OpenAI and Anthropic came back with zero
- Quasar 438B: Europe's Leading AI Model
- v2.1.258
- v2.1.257
- v1.18.26
- How AI-native companies turn workflows into operating capability
- How law firm Gilbert + Tobin governs and scales AI with OpenAI
- The latest AI news we announced in August 2026
- Mapping global methane emissions from space with deep learning
- Real-Time Intelligence with IBM Time Series Models on Confluent
- BenchMIRT: What are LLM benchmarks actually measuring?
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- How we make AI coding more cost efficient without sacrificing task quality
- How we could save petabytes of cache storage with Zstandard and Pingora
- TeamCity 2026.2: Pipelines General Availability, BYOK for AI Assistant, and More
- The MPS 2026.2 Early Access Program Has Started
- Stop Guessing at Hard Faults
- How to Handle Errors in Go
- The Grails Plugin Has a New Home: Apache Grails
- Ensuring Code Compliance in Public Sector Software Projects
- Authenticating TeamCity Builds to External Services With OIDC
- Building Reproducible AI Evaluation Workflows with Docker Sandboxes
- Below the Harness: Governing a Multi-Model, Multi-Harness World
- The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough
- Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
- Building an Adaptive Agentic Cybersecurity System with NVIDIA Nemotron
- Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA
- Facilitating AI integration with simplicity at scale
- Trump may be forced to reveal secret rules feds use for AI safety testing
- Wonderful more than doubles its valuation to $5B in under 6 months
- India’s richest man now wants to turn aging computers into AI-ready PCs
- HiddenLayer nabs $100M as enterprises rush to secure their AI deployments
- Adobe acquires Indian market intelligence startup Rilo
- AfterQuery reportedly becomes Y Combinator’s fastest-ever unicorn, now valued at $3.2B
- Google’s Android update tackles motion sickness, accessibility, and more
- Sequoia-incubated Empirik launches with $21M to predict outages before they happen
- Amazon Alexa can now alert you when something new might tempt you to shop
- AIR raises $50M to help companies vet the skills and add-ons AI agents use
- Fambot introduces an ‘AI chief of staff’ for families
- Apple shares ‘shocking evidence’ against former employee accused of stealing company data for OpenAI
- Google is sending MrBeast into the wilderness, armed with AI
- NYC bans AI use for students until they reach high school
- Google needs Hollywood more than the studios need AI
- OpenAI delayed its new model’s development after the Hugging Face hack
- The rise of AI ‘civilizations’ and the fall of corporate responsibility
- Swiggy Uses 350+ Features and Multi-Task MLP to Predict Customer Lifetime Value
- Presentation: Beyond Prompting: Context Engineering for Production-Grade AI
- Cloudflare Adds Optional OAuth Scopes, Letting Developers Mark What Users May Decline
- OpenClaw 2.0 Releases with Simplified Setup and Collaborative Agents
- InfoQ previews the September cohorts of its online certification programs
- HCP Terraform Positions Itself as the Control Plane for AI-Driven Infrastructure
- AI Efficiency Could Cost Us the Next Generation of Experts
- Cash In on the AI Boom by Renting Out Your Spare Compute
- Protests against AI data centers play into China's hands, Trump says
- Google Deepmind's new chief says frontier AI leadership is the only thing that matters
- HyperWorld: Hypergraph-Structured State Serialization Improves Learned Textual World Models
- I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
- Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
- Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
- Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
- OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
- SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces
- UI-Venus-2 Technical Report
- EULER: Exploring Underused Links with Evidence-Checked Return for Multi-Agent Mathematical Discovery
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
- MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts
- AI Morbidity and Mortality: A Framework for Clinical AI Failure Review
- Different representation learning objectives recover distinct latent structures from the same psychometric data
- Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations
- Recursive Criticality of AI Self-Improvement
- IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
- Asymmetries in Spontaneous and Instructed Deception
- LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
- ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
- AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning
- ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback
- Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
- Invalidation Contracts for Cross-Episode Agent Memory
- Authority Bias in Conversational Search Engines for Academic Paper Recommendation
- Hypotheses-Guided Self Distillation for Continual Personalization
- The Answer Is Not the Argument
- Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching
- The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems
- The Assistant's Ideal Self
- Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective
- SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning
- Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
- A Stable Aggregation Method for Quantum Federated Learning
- Dr. Claw: An AI Scientist Workspace for Vibe Research
- RestoreBench: Can AI Agents Restore Power Flow Convergence?
- Dependency-Aware Chain-of-Thought Compression for Financial Reasoning
- SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
- SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents
- Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations
- mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers
- Towards a Belief-Based World Model for LLM Agents
- EGT-KG: Evidence-Grounded Typed KG Retrieval for Practical Scientific QA with Small Language Models
- The Privacy-Hallucination Tradeoff in Differentially Private Language Models
- Validity-Aware Jailbreak Evaluation for Large Language Models
- Wave Function Backpropagation with Explicit Temporal-Interval Dynamics
- CoVer: Conflict-Aware Claim Verification
- When the Algorithm Becomes the Brand Crisis: A Sociotechnical Theory of Distributed Responsibility and Accountable Transparency
- ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation
- Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation
- VoiceLongMemEval: Do Assistants Remember How You Sounded?
- Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs
- Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random
- Same Request, Different Boundary: Evaluating Cybersecurity Assistance across Conversational Contexts
- Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing
- Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs
- REVISE: Validity-Guided Recovery for Online Revisions in Agent Workflows
- DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation
- Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search
- SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task
- Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets
- Triple-Bottom-Line Sustainability of Language Models for Edge AI: A Comparison Between SLMs and Quantized LLMs
- Value Over Language Model: Detecting Original Contribution in Writing
- ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything
- A Closed-Loop Evaluation of Capability Loss and Recovery in Compressed Driving Policies
- SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification
- Agentic Empirical Asset Pricing: Methodological Foundations
- Escaping Redundant Reasoning: Structure-Aware Search for Inference-Time LLMs
- ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents
- S^3martCirc: Self-supervised Smart Circuit Discovery
- Automated Tree Knowledge Graph Construction using Ontology Expansion and Retrieval from Vietnamese History Textbooks
- DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
- When Features Become Instances: Inverted Contrastive Learning for Unsupervised Feature Selection
- StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
- Towards a Reliable and Practical Eval Pipeline
- One Policy, Any Budget: Internalizing Budget-Aware Search via Reinforcement Learning
- AnalysisBank: An Expert Analysis Pattern Library for Financial Report Generation
- Polished but Unresolved: Identifying Late-Stage Pressure States in Long-Horizon Tool-Use Agents
- FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation
- Towards Generalizable Visually Grounded Exploration of Household Devices
- Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources
- Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems
- Beyond the Clock: Measuring the Value of Adaptive Revision
- FractalNet-Based Heterogeneous Federated Learning for Orbital Edge Intelligence in Satellite Mega-Constellations: A Wildfire Case Study
- Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning
- Denoising Diffusion Generative Models Secretly Calculate Attentions
- CacheBridge: Efficient Cross-Model KV Cache Transfer
- CARE: Contrastive Anchor-based Rubric Evolution for Large Language Model Post-Training
- In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?
- RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation
- VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences
- Few-Shot Out of Domain Intent Detection with Covariance Corrected Mahalanobis Distance
- CoBRA: Learning Tool-Use Boundaries via Counterfactual Margins
- Figures as Programs: Recursive Generation of Editable Scientific Figures
- Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees
- Data-Driven Persona-Conditioned Agents for A/B Test Simulation
- AgentFactory: Towards Automated Agentic System Design and Optimization
- QILP-0: Constructing Observational Declarative Twins of Quantum Circuits
- WorldBench: Culturally Grounded Benchmark for Multilingual Agents
- User Representation via Cross Multi-source Behavior Pre-training for Mobile Games
- ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning
- Space Generative AI with Solar Energy Harvesting
- Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs
- Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate
- FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue
- H2Table: Hierarchical Hypergraph-Enhanced Large Language Models for Complex Table Reasoning
- Prompt-Robust Language Models: Which Training Strategies Work?
- Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
- Dual Process Motion Planning
- Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents
- Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems
- A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
- Automated Event Log Generation from Unstructured Text Using Finetuned LLMs
- LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting
- Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
- SymFold: Synergizing Evolutionary and Structural Priors for Accurate Protein Inverse Folding
- EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems
- Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations
- EdiTikZ: Scientific Figure Editing from Revision Trajectories
- Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
- When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation
- EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
- Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
- Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers
- InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed Information
- Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
- trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
- RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery
- Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
- REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
- Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness
- From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
- Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
- ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation
- CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
- DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
- A Formal Analysis of Agent Payment Protocols
- ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
- RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
- Medical Causal Hypothesis Verification with Large Language Models
- Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
- Scientific Agent Skills: A Library of Procedural Knowledge for Research Agents
- OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
- Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
- Life Operators: a self-evolving framework for multiscale life modelling
- Auditing Harness Tampering in Self-Improving Agents
- AutoXRD: Autonomous LLM Agents and Comprehensive Evaluation for Powder Diffraction Analysis
- RW-LoRA: Communication-Efficient Decentralized LoRA Fine-Tuning via Random Walks
- KItCAT: Knowledge Injection via Input Corruption for Auto-regressive Training
- Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
- Commit-first LLM judging inherits the judge's own errors
- Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence
- Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
- Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
- Flawed in Nature, Perfect through Evolution
- Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
- Do General NLP Embeddings Capture Ontological Reasoning?
- Intelligent Edge Computing
- Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
- Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost
- WHALE: A Simple Recipe for Joint Harness-Weight Optimization
- Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
- Rock, Paper, Scissors, ... Dynamite - A Model of Disruption from New Technologies
- QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
- Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes
- CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
- CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
- Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems
- Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection
- WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization
- Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains
- Workload Identification with Physical Side Channels for AI Governance
- A Human-AI Theorem Connecting Spontaneous and Field-Induced Mechanisms of Collective Behavior in One Dimension
- The Curse of Multilinguality in Lexical Normalization
- Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
- Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces
- Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
- Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
- Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning
- Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
- FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
- Risk-Aware Decision-Making for Autonomous Overtaking: A World Model-Based Mixture-of-Experts Framework
- (V)LMs generalize beyond surface co-occurrence: Evidence from cross-modal number agreement
- Capability-Gated Language Models: Security Composes, Utility Does Not
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
- HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference
- Does Reasoning Mitigate Backdoor Attacks? A Neuro-Symbolic Perspective
- Operational Regimes in Non-Convex Optimization: A Multiplier-Based Taxonomy
- Higher Structures in Deep Learning
- Exploring Collaboration between a language and a non-language agent
- EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities
- Beyond Token Positions: Safety Alignment Across Denoising Steps in Diffusion Language Models
- Independent Reinforcement Learning in Discounted Markov Games
- RecalibrateGPT: AI Fatigue Resilient Conversational Interfaces
- The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
- The Safeguard Worked. Is the LLM System Safer?
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
- Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
- EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
- EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection
- WiseSpec: Requirements-Driven Agents for Code Generation
- A Mathematical Framework for Legacy, Governance, and Decision Integrity in Enterprise AI
- GeoPAR: Large-Scale Multi-Agent Combinatorial Optimization with Geometry-Guided Parallel Autoregressive Learning
- Predicting Program Exit Code with LLMs and Programming Language Semantics
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
- Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning
- Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search
- Restrict, Don't Retrain: Inference-Time VLM Guidance for Zero-Shot Aerial Segmentation
- Breaking the Structural Identity: Personalized Federated LoRA Fine-tuning under Rank Heterogeneity
- TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data
- EEG-AS: Instance-Level Foundation Model Selection for EEG Foundation Models via Behavior Reconstruction
- Visual Framing for News Stance Detection via Image Generation
- A Study of Hidden-State Optimization Order in Predictive Coding Networks
- Differentially Private Paired Table-Image Multimodal Synthesis
- Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
- Are You Thinking What I am Thinking? : Examining Conceptual Separation in Neural Architectures
- VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM
- Solaris: Towards Interfaces That Are Generated, Not Coded
- Instella-MoE Technical Report
- MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries
- Agentic programs: an emerging form of scientific software in computational materials science
- Ctrl-F-Resist. Practices, Challenges, and Technical Needs of Civil Society Organizations Monitoring the Far-Right Online
- HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution
- Visual Attention Faithfulness in Vision-Language Models is Heterogeneous
- Replacing Training with Memory: Listwise Selection for Text-to-SQL
- Probabilistic Model Checking of Autoregressive Neural Sequence Models
- A Checklist to assess the energy and carbon impacts of ML/AI applications in Earth System Modeling
- ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection
- Does Fault Localization Beat a Fresh Attempt? A Placebo-Controlled Study of Test-Guided Code Repair
- Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation
- Vision-Language-Guided Pseudo-Labels for Unsupervised Domain Adaptation in Semantic Segmentation for Waste Sorting
- Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking
- Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO
- DualStake: Dual-Path Confidence Calibration in Deep Research Agents
- Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches
- From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding
- Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
- The zbMATH Open Knowledge Graph: Tracing Centuries of Mathematical Research
- Disclosure-Gated User Simulation for Companion-Agent Evaluation
- Semi-Supervised Virtual Staining via Morphology Preservation and Histopathological Realism Constraints
- On the Human and Computer Alignment of Attribute-Based Music Matches
- Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages
- Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close
- SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models
- A Network Science Perspective on Evaluating Deep Graph Generative Models
- On Synthesis of Metric Interval Temporal Logics
- Causal Evidentiary Governance for High-Risk Machine Learning Systems
- ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives
- From Truncation to Commitment: Persistent Context in Uniform Discrete Diffusion
- HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
- Lagged Coupling: Internal Representations Become Readable Before They Become Causal
- Text-guided flow matching enables sample-efficient crystal structure generation
- StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions
- Hints Help But Do They Teach? Evaluating Skills Transfer in Code Generation
- EDRAC: Benchmarking Arabic Dialect Reading Comprehension
- DNC-IMM: Early Lane-Change Intention Recognition via Neural Calibration Based on Driving Context Information
- Revisiting Face Recognition for Monozygotic Twins: The Celeb Twins Test Set
- StainPresetNet: Stain Preset Network for Fast Multi-to-Multi Stain Normalization
- Superposed Latent Autoencoder
- Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion
- Towards AI-Assisted Clinical Trial Matching: Practical Considerations, Multicenter Evaluation, and Real-World Deployment
- Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening
- Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges
- REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
- Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling
- MutMem-V2: Cryptographically Authorized Mutation in Persistent Agent Memory Portable Verification and Reproducible Evidence
- From Language to Behavior: Scaling Sequence Transformers for Industrial Recommendation Ranking with Rec-Native Designs
- Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents
- One Prompt Is Enough: Watermark Laundering Through Foundation Image Models
- The Constitutional Coverage Trilemma in AI Governance
- TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models
- Some Emotions Run Deeper: Layer-wise Probing and Causal Intervention in Large Language Models
- EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
- HiLRP: Toward One Trustworthy Explanation for Vision Transformer: Conservation-Valid Attribution via Attention Primitives
- GazeRefine: Expert Gaze as a Test-Time Prompt for Training-Free Medical Image Segmentation
- MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval
- Bandits in Prod: Hyperparameter Optimization at Inference Time
- Probing Factual Knowledge Transfer with Training Data Interventions
- Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs
- CHARM: Character Hallucination for Multicultural Role Play Benchmark
- PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction
- Measuring consistency via ensemble margin and local prediction variability: Auditing decision systems in the presence of predictive multiplicity
- Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
- Provably Safe Sim-to-Real Transfer
- Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading
- Learning Sparse Decision Trees via Transformer Variational Auto-Encoders
- Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search
- When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
- Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents
- GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
- Rethinking Learnability in Offline Data-driven Optimization
- Optimizing Byzantine Node Placement in Decentralized Federated Learning
- LatentPress: Context Compression Beyond Text and Vision
- TempCloze: Can Video-LLMs Identify the Missing Middle?
- Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data
- Can LLMs Design Video Coding Tools? A Case Study on Planar Mode
- A Mathematical Theory of Reusable Neural Bases for Network Compression
- BS: Take the Hint - Interactive Multitracer PET/CT Lesion Segmentation with a Scribble-Conditioned ResEnc U-Net
- Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
- H3-World: Turning Language Understanding into World Control
- From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification
- Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
- Designing Proactive Thought Partners for Writing
- Mechanism Design for Alignment and Control
- The Rise of Verbal Reinforcement Learning
- CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?
- Adaptive Critical Token-Aware Retrieval for Repository-Level Code Generation
- Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
- GeoGR^2:Zero-Shot Geospatial Inference via Geostatistically-Guided Iterative Refinement with LLMs
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs
- Compositional Machine Design as Program Synthesis with LLMs
- HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
- LifeAgentBench: Benchmarking LLMs for Long-Horizon, Cross-Dimensional Lifestyle Health Reasoning
- Think Like a Doctor: Conversational Diagnosis through the Exploration of Diagnostic Knowledge Graphs
- MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
- Ontology-Guided Neuro-Symbolic Inference: Grounding Language Models with Mathematical Domain Knowledge
- HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
- TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning
- D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
- Causal Probing for Internal Visual Representations in Multimodal Large Language Models
- SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
- UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities
- Diffusion Large Language Models for Visual Speech Recognition
- The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic
- DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reasoning Models
- Flow Reasoning Models: Turning Flows Into Efficient Recurrent Reasoners
- Self-Evolving World Models for LLM Agent Planning
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
- AREX: Towards a Recursively Self-Improving Agent for Deep Research
- UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
- MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
- Reasoning-supported Robustness Validation of Automotive E/E Components
- The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
- Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks
- Electronic Navigational Chart Change Classification
- RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
- The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning
- pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning
- AI Alignment through a Game-theoretic Lens: A Survey
- AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative Investment
- Automated Researchers Can Reliably Mitigate Alignment Failures
- Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling
- Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food Knowledge
- Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling
- Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
- Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
- Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- Building Expressive and Tractable Probabilistic Generative Models: A Review
- FedReview: Review and Dispose Poisoned Updates without Validation Datasets or Historic Knowledge
- Keep Everyone Happy: Online Fair Division of Numerous Items with Few Copies
- Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models
- X-SG$^2$S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks
- SARTM: Segment Any RGB Thermal Model with Language aided Distillation
- A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
- Towards Provable and Scalable Training of Quantized Neural Networks with Ising Optimization
- ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation
- Unsupervised Partner Design Enables Robust Ad-hoc Teamwork
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Injection
- SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance
- TopoAlign: A Framework for Aligning Code to Math via Topological Decomposition
- One-shot Style Transfer LLM log-probabilities for Authorship Attribution and Verification
- Taming Modality Entanglement in Continual Audio-Visual Segmentation
- Can machines think efficiently?
- Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
- Individualized Algorithmic Advice as a Strategic Signal on Competitive Markets
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
- A Machine Learning-Driven Solution for Denoising Inertial Confinement Fusion Images
- The Alexander-Hirschowitz theorem for neurovarieties
- 3D-Consistent Multi-View Editing by Correspondence Guidance
- Multilingual Medical Reasoning for Question Answering with Large Language Models
- Hidden State Poisoning Attacks against Mamba-based Language Models
- A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning
- Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
- FloydNet: A Learning Paradigm for Global Relational Reasoning
- Persistent Entropy as a Detector of Phase Transitions
- Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning
- Make Some Noise: Unsupervised Remote Sensing Change Detection Using Latent Space Perturbations
- Channel-Adaptive Edge AI: Maximizing Inference Throughput by Adapting Computational Complexity to Channel States
- MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery
- Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
- Guided Prompt Evolution for Vision-Language Models Adaptation
- RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction
- Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation
- V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
- SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction
- MineDraft: A Framework for Batch Parallel Speculative Decoding
- Revealing Multi-View Hallucination in Large Vision-Language Models
- APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay
- VectorGym: A Multi-Task Benchmark for SVG Code Generation, Sketching and Editing
- Oblivion: Self-Adaptive Agentic Memory Control through Decay-Driven Activation
- IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
- Training-Free Refinement of Flow Matching with Divergence-based Sampling
- DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
- KV Cache Offloading for Context-Intensive Tasks
- What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
- Why Fine-Tuning Encourages Hallucinations and How to Fix It
- Global Attention with Linear Complexity for Exascale Generative Data Assimilation in Earth System Prediction
- Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
- The Topological Trouble With Transformers
- Universal Approximation of Nonlinear Operators and Their Derivatives
- Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
- When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer Selection
- MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs
- Skill Reuse as Compression in Agentic RL
- PlanarBench: Evaluating LLM Spatial Reasoning via Planar Graph Drawing
- Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
- GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
- Enabling KV Caching of Shared Prefix for Diffusion Language Models
- DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment
- Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS
- Generativism: Toward a Learning Theory for the Age of Generative Artificial Intelligence
- AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
- ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
- Steer, Don't Solve: Training Small Critic Models for Large Code Agents
- Energy-Based Transformers as Predictors of Reading Difficulty
- Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology
- Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?
- DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
- LUNA: Learning Universal 3D Human Animation Beyond Skinning
- Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
- Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
- How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
- Backspace as a Natural Experiment: An Accelerated Failure Time Model of Selective Post-Error Motor Impairment in Parkinsons Disease
- Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
- Can We Trust In-Distribution Success? Locked Evaluation Reveals Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation Prediction
- When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation
- Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
- Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
- Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements
- Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight
- HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
- Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
- Performative Privacy: When Differential Privacy Maximizes Utility
- Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
- REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling
- Arkios: An Open Bilingual English-Nepali Language Model Trained From Scratch, with a Devanagari-Aware Tokenizer
- E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval
- Cost-efficient Active Learning for Referring Image Segmentation and Grounding
- BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs
- Calibrating Small Language Models for Claim Check-Worthiness Detection
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Stochastic complexity of vectors containing cluster structure
- Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
- Safin-1: Safety from Within through Memory-Native State Evolution
- Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification
- Generative artificial intelligence for reliable mechanistic reasoning for corrosion
- Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization
- Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment
- Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance
- Neural means and kernel corrections for operator learning
- A Multi-Branch Feature Fusion Approach for Health Misinformation Detection and Propagation
- How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks
- Group Adaptive Clipping Policy Optimization
- CRAD: Class-wise Reliability-Aware Distillation for Decentralized Heterogeneous Federated Learning
- Can LLMs Use Relational Transformer Embeddings?
- Context Window Failures in Relational Foundation Models
- AdaptNTK: Adaptive Uncertainty Quantification and Active Learning for Neural Network Potentials
- A hybrid quantum-classical neural network for learning to route
- VATO: A Vortex-Force-Aware Transformer Operator for Unsteady Separated Aerofoil Flows
- Learning Task-Specific Antibody Representations via Function-Aware Masking
- Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials
- DeSyR: A Decoupled Symbolic Recovery Framework with PINN-Guided Structure Search and Physics-Informed Coefficient Refinement
- GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting
- Manifold-Aware General Coded Computing for Straggler-Resilient Distributed Computing
- CRAFT: Fine-Tuning Pre-hoc Explainability in AI-native 6G RAN
- Topological Steering
- DK-GBMKKM: Dynamic Kernel-Space Granular-Ball Multiple Kernel $k$-Means Clustering
- HarmoCore: Functional Latent Diffusion for Sparse Reconstruction of Oscillatory Wave Fields
- Verdict Instability of OOD Scores under Reference Resampling
- MUGEN: Generating Unlearnable Graph Examples for Multiple Learning Tasks
- Patterning in Practice: Debiasing Reward Models with Susceptibilities
- Online Self-Weighted Fine-Tuning
- Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis
- How Do Language Models Choose Between Context and Memory?
- Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation
- Subspace Levenberg Marquardt Algorithms in Training Neural Networks
- Conditional Flow Matching for ML-Based Inverse Design Problems
- MemoryWalker: Stop Training Agents on Contexts They Never Saw
- iPINN for Broadband CARS Phase Retrieval: A Framework for Function Approximation and Inverse Modeling Problems in Nonlinear Spectroscopy
- Poisson-Gamma Dynamical Systems with Time-varying Transition Dynamics
- When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting
- The Multiple Timescales of Gradient Descent on the Edge of Stability: A Perturbative Derivation of the Central Flow
- SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations
- Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
- Modelpedia: A Catalog of Model Findings for the Meta-Science of AI
- Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
- Neural Symbollic Regression Using Deep Learning and Sparse Modelling
- Replicating TRACE: A Practitioner's Guide to Its Threshold and Particle Budget
- When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting
- Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras
- CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs
- Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training
- Births are difficult to predict even with rich survey and full-population register data
- Recent Developments in Transformer Inference Deployment on FPGA Platforms: A Survey
- Multi-Head Self Attention is a Parameter Identification Mechanism
- Post-Training Science for Supervised Fine-Tuning
- Solving In-Table Prediction Problems by Deep Neural Networks with Performance Evaluation Using Synthetic Data
- Position: Privacy Is a Claim, Not a Property of Synthetic Data
- One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- Contribution-Aware Bandwidth Allocation for Multimodal Split Learning
- Predicting Subsurface Abnormalities Growth using Physics-Informed Neural Networks
- CATeye: Coupled Attribute-Topology Invariance Learning for Voucher Abuse Detection
- TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution
- Edge-Girth as a Structural Edge Feature for Graph Neural Networks
- Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning
- Quantum Sparse Autoencoders for Q-Matrix Estimation in Cognitive Diagnosis
- NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games
- Gradient-Update Mismatch: Rethinking Conflict-Free Training of Physics-Informed Neural Networks
- The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally
- ES-AHD: An Evolution Strategy Framework for Automatic Heuristic Design
- Dense Weak Hiding: Closing Complexity Gaps in Nonconvex and PL Finite-Sum Optimization under Individual Smoothness
- AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes
- Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
- Exact Global MCMC with Denoising Diffusion
- Lightweight Adaptation of EEG Foundation Models for Stroke Motor Imagery Decoding: Domain Shift and Subject-Level Robustness
- TRUST: Threshold-Recalibrated Uncertainty-Safe Training for Certified Dismissal in Breast Cancer Screening
- Towards unsupervised representation learning for quantum data: quantum models with inference and generation
- Hidden relationships in a document-derived property graph: top-k chunk embeddings and inverse-distance weighting over a dynamically evolving ontology
- NeuroPriv: Adversarial Representation Learning for Privacy in Wearable EEG Systems
- A convolutional framework for detecting event-driven dynamics in energy price series
- DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
- Accelerating Chemical Kinetics for Exoplanet Atmospheres using Neural Networks
- Physiological Information Reliability: Cross-Layer Adaptive Resource Allocation for Cardiovascular Sensing
- Fractal dimension predicts quantum kernel collapse in angle-encoded data
- Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
- Soft-Argmax for the Projective Plane via the Veronese Embedding
- Real-Time Neuromorphic Spectrum Intelligence Simulator
- BeamRMX: Radiation-Pattern-Driven Learning for Generalizable Beam Radio Map Prediction and Beam Management
- Disciplined Bilevel Programming
- Controllable Image Captioning with Prompt-Conditioned Scene Rewards
- Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption
- MaskCode: Mask Transformer for Feedback-Assisted Coding With Linear Block Codes
- Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models
- Dense Process Supervision for Search Agents via Fact Utility Estimation
- The Visual Insensitivity Gap: Diagnosing When Vision-Language Models Fail to Use Visual Evidence
- Sharp Mixed Spectral Barron Regularity of Coulombic Many-Electron Wave Functions
- Direct Optimization of a 3D Finite-Source Reflector via Neural-Network Parameterization
- Web Price Extraction: State of the Art and an Adaptive Browserless Implementation
- Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC
- Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models
- Relational Task Generation Language: A Declarative Specification Framework for Relational Deep Learning
- Matched Queries for Curvature and Density at Branching Junctions
- Exploring Sparse Autoencoders in Text-Based Causal Confounding Adjustment
- mzCache: On-Device LLM Memory Management under Multitasking
- Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
- Exact Risk-Complexity Laws for Projective Boundaries in Scenario Optimization and Distribution-Free Certification
- Investigating Linear Probe Robustness to Linguistic Register, Medical Specialty, and Corpus Shifts in Medical QA
- On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
- Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds
- Sierpi\'nski--Knopp Wasserstein Distance for Persistence Diagrams and Applications to 2-Wasserstein Approximation
- Variable Selection for Feature-Based Newsvendor
- Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation
- Beyond Scores: Understanding LLM-as-a-Judge Mechanisms in Summarization Evaluation
- QABBA: Symbolic Time-Series Compression via Integer-Quantized Aggregation
- Multi-View Causal Discovery without Non-Gaussianity: Identifiability and Algorithms
- Efficient Learning of Balanced Signed Graphs via Sparse Linear Programming
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
- Recurrent State Encoders for Efficient Neural Combinatorial Optimization
- A Compositional Kernel Model for Feature Learning
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective
- Iterative GRPO: Batch-Online Policy Iteration for Multi-Turn RL via Single-Turn RLHF
- Freeze, Diffuse, Decode: Geometry-Aware Adaptation of Pretrained Transformer Embeddings for Antimicrobial Peptide Design
- Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs
- Control Variate Score Matching for Diffusion Models
- Variance-Adaptive Muon: Pre-Orthogonalization Variance Modulation for Efficient Language Model Pretraining
- Breaking the Reasoning Horizon in Entity Alignment Foundation Models
- Efficient Adaptation of ROMs for Unsteady Flows Using Data Assimilation
- Inverse Reconstruction of Shock Time Series from Shock Response Spectrum Curves using Machine Learning
- Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
- Rigorous Error Certification for Neural PDE Solvers: From Empirical Residuals to Solution Guarantees
- Process-Aware AI for Rainfall-Runoff Modeling: A Mass-Conserving Neural Framework with Hydrological Process Constraints
- Reparameterization through Coverings and Topological Weight Priors
- Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
- LLM-driven design of physics-constrained constitutive models: two agents are better than one
- Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
- PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning
- What Do Students Learn? A Feature-Level Analysis of Dark Knowledge
- RECAP: Regression Evaluation for Continual Adaptation of Prompts
- Learning to Refine Hidden States for Reliable LLM Reasoning
- Beyond AHI: An Interpretable Causal-Discovery-Guided Framework for Sleep Recovery in Connected Health
- Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
- Shallower ReLU Network Representations via Exact Linear Algebra
- S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
- Reading the Gate, Not the Interference: Output-Side Interference Measurement Does Not Track Merge Collapse
- Non-Parametric Spatiotemporal Trajectory Prediction via State-Conditioned Transition Sampling
- Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers
- Stress Testing Unlearning Algorithms
- The Frame Kernel Method for Multiscale Operator Learning
- Deep learning based numerical approximation algorithms for stochastic partial differential equations
- GENIE: Watermarking Graph Neural Networks for Link Prediction
- Generalization Bounds for Markov Algorithms through Entropy Flow Computations
- Online simultaneous inference for quantiles via smoothed stochastic gradient descent
- On the Existence of Consistent Adversarial Attacks in High-Dimensional Linear Classification
- FlexP-SFT: A Flexible Aggregation-Free Framework for On-Device Personalized Split Federated Fine-Tuning of LLMs
- Integrated Noise and Safety Management in UAM via A Unified Reinforcement Learning Framework
- Fair Minimum Labeling: Efficient Temporal Network Activations for Reachability and Equity
- Silence is Golden: Mitigating Hallucinations in Large Audio-Language Models via Layer-Weighted Vector Steering
- Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
- Probabilistic Multi-Agent Aircraft Landing Time Prediction
- CADKnitter: Compositional CAD Generation from Text and Geometry Guidance
- Modeling Information Blackouts in Missing Not-At-Random Time Series Data
- DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
- Auditing Frozen-Encoder Anomaly Detection Across Mechanical Systems: Representation Provenance, Calibration, and Protocol Effects
- Online Regime-aware Calibration for Black-box Social Simulators via Posterior-assisted Evolutionary Dynamic Optimization
- Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging
- Is Knowledge Distillation Actually Greener? A Case Study in Machine Translation
- A penalised Saito functional for heuristic search of free line arrangements
- FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet
- Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding
- Polarizable atomic multipoles for learning long-range electrostatics
- DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
- Three-dimensional Conditional Diffusion Models for Cosmological 21 cm Lightcone Emulation
- HiMPO: Hindsight-Informed Memory Policy Optimization for Less-Entangled Credit in Long-Horizon Agents
- Explicit Interaction Architectures for Dynamical Learning: A Controlled Study of Structural Inductive Bias
- Closing the Operational Gap in Semantic Caching
- EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography
- A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification
- Field-Aware Agent Skill Retrieval
- Coordinate-Residual Physics-Driven Neural Network for Inverse Scattering Imaging
- Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms
- ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents
- MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
- Common-Center Geometry and Certified Radial Reconstruction for Energy-Form Full Conformal Regions
- TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization
- TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories
- Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents -- A Source-Code Study of Eleven Systems
- Structure-Behavior Coalescence and the Limits of Traditional Systems Theory
- What Is a System? An Interaction-Based Account of Structure-Behavior Coalescence in General Systems Theory
- Can MCP Clients Decide What to Do After Failure? A Result-Only Actionability Audit
- Framework and Benchmark for Code-Driven Agentic Testing in Web Development
- Empirical Software Engineering in Practice: Insights from Google
- Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork
- Exploring Quantum Software Testing Across Research and Practice: Emerging Results from a Multivocal Literature Review
- Revisiting Feedback-Driven LLM Code Repair: A Replication and Exploratory Java Extension
- Audit-First Rollback Semantics for Safety-Critical Deployment Pipelines
- What Survives the Next Model? Benchmarking LLM-Based Techniques Against Single-Prompts
- Fine-Tuning Large Language Models to Classify Pull Request-Issue Alignments: Going Beyond Prompting
- Reliable LLM-Generated Programs for High-Energy Physics Experiments through Graph-Grounded Software Knowledge
- Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance
- What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- The Data Problem in Software Vulnerability Analysis: Artifacts, Quality, and Consumption
- RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
- SilentProbe: Measuring Silent Failure in Production APIs Used as Agent Tools
- Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
- Beyond Locks and Thread IDs: Static Data Race Detection Off The Beaten Path (Extended Version)
- Bounded, Indeterminate, or a Bug: A Condition-Aware Oracle for Differential Testing of SQL Aggregates
- Operation-Type-Aware Client Routing for Leader-Based Consensus Datastores
- Federated Trust for Embodied Robot Capability Marketplaces
- Smart Contracts Claimed Vulnerable by the CVE Database, with Labels and Source Locations
- The Popularity Hypothesis in Software Security: A Large-Scale Replication with PHP Packages
- Essence Coach: A Bot for Software Practice Adoption
- Oops!... I did it again. Analysing and Handling Conclusion (In-)Stability in Socio-Technical Software Engineering
- Measuring Computer Science Enthusiasm: A Questionnaire-Based Analysis of Age and Gender Effects on Students' Interest
- From Rookie to Pro: Social Engineering LLMs for Automated Vulnerability Exploitation in Enterprise Software
- Test vs Mutant: Adversarial LLM Agents for Robust Unit Test Generation
- ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance
- SEDCoT: Enhancing LLM-Based COBOL Code Translation via Symbolic Execution and Delta Debugging
- Automatic Model-Hardware Co-Adaptation for Heterogeneous AI Accelerators
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
- Normalized Fascism in Open Source: $12 Million Given to DHH
- The load-bearing vocabulary of Claude
- The Endless Temptation of Claude
- Is Minifying CSS Necessary? (2023)
- A Crash Course in Predicate Logic
- Dependent if expressions without dependent types
- Fine, I’ll build my own text editor
- What will you do after tech?
- GentleOS/16 hobby OS for vintage 16-bit PCs
- Static Allocation, Constant Work
- TinyGo 0.42 - Recover Is Real
- ESP32 as counter-surveillance platform
- Read your own writes, off the primary
- Revisiting Joel's Test - exe.dev blog
- Implementing FMA and finding bugs in C and Rust standard libraries
- Bluefin is a capability system
- Zuzai, a new word, indicates the absence of AI
- New things for regular expressions in PostgreSQL (pg_tre and pg_re2)
- humanity has built the records of FATE by accident
- A bicycle for the mind
- "iT woRKs BeTter in THe aPp!!"
- Why I'm excited about effect systems (2025)
- BizNode gives you a full web dashboard at localhost:7777 — manage leads, conversations, knowledge base, and settings in one...
- Content Moderation That Never Sees Your Private Messages
- How to Build an AI Moats That Actually Last
- We’re Building Websites Backwards: An Information-First Architecture for the AI Web
- AI APIs for Crypto Trading Signals - Complete Guide
- Free Unlimited Web Search for AI Agents
- The Website Is No Longer the Center of Commerce
- Revenue Strategies for AI API Services
- MCP vs RAG Explained (Complete Beginner Guide with Architecture)
- My Take After a Year of Using Cursor — 62-Day Longest Streak
- Beyond DOM Scraping: Building "THE LAST TERMINAL" with WebMCP
- AI reshaping how we build, review, and trust code
- The Model Offered a One-Liner. I Spent 48 Hours Refusing to Run It.
- The Prompt Is Not a Lockfile: A Provenance FAQ
- ADLC: The Lifecycle Taking Shape
- The Memory Limit That Didn't Kill Anything
- Qwen Code 400 "failed to parse grammar" against llama.cpp? Here's the fix
- Using LLMs for Crypto Market Analysis in 2026
- Make Your AI Sound Truly Human: Discover `avoid-ai-writing`!
- You can't delegate what you can't verify
- My AI Gateway Added 400ms to Every Request. Here's Where It Went
- Gemini Advanced review: 1M context window changes everything in 2026
- JSON Mode Makes Your LLM Dumber: The Constrained Decoding Trap
- Your System Prompt Has a Shelf Life: Maintaining Prompts as Models Improve
- Fable 5.1 Max gave me the most reasonable local setup guide
- Fable 5.1 made a Minecraft mod for $20
- Differences Between Fable 5 and Fable 5.1 on MineBench
- Well I almost got prompt injected
- Anthropic, we want Fable back into the pro plan!!!
- Fable 5.1 is insane and it burned usage, which is fine. Anthropic just needs to nail Opus 5.1
- First one to out-lead Claude on Code Arena in a long time. Also first Chinese ever. Landscape is changing
- Fable 5.1 - Reminder re Watermarking and Note for attorneys
- Anthropic really doesn’t seem to value its $20 subscribers anymore
- Claude Code can now build a working internal tool in minutes without writing any application code. We put ToolJet behind an MCP server (50 small tools, MIT).
- Is anyone else burning through tokens way faster with Fable 5.1?
- Gone in 60 seconds
- Does anyone feel like Fable 5.1 has been nerfed since release?
- I asked Claude to draw itself after analyzing our chat history.
- Fable 5.1 is finally a understandable model
- Anthropic's Fable 5.1 Guide on dense prose is dense Claudish slop
- The part of Claude Code that concerns me most is how easy it is to approve work I only half understand
- Fable 5.1 is out it’s amazing — it’s terrible — they nerfed it —
- (1st impressions) I burned an entire Claude Max 20x on Fable 5.1 in 8 hours
- Is Claude Fable 5.1 spawning 126 agents for simple tasks for anyone else? It burned ~8.4M tokens on a 5-file i18n audit!
- Opus 5 make me laugh for the first time in months
- Weekly Thread: Project Display
- If you believe a 19 year old makes $300k a month from an AI agency you deserve to get scammed by his course.
- The better local models get, the harder it is to justify buying a box to run them on.
- What is one thing AI still makes harder than it should?
- How do you know your long shared prefix is really being cached?
- AI sales reps failed because they were trained on sh*tty data.
- What makes an AI agent actually useful?
- What should I know before implementing AI agent security?
- First real project I've built, a multi-agent "personal executive AI" instead of one big assistant. Would love feedback.
- anyone has experience with proactive ai agents?
- Everyone says write evals for your agent. But what should you actually test?
- What’s the most “agentic” thing you’ve built that actually survived contact with real life?
- When you went from the free tier to paying on an AI tool, what was the exact thing that made you do it?
- Is a fully automatic, cheap, well-working, outreach agent possible?
- Would you actually use an “infrastructure layer” for your AI agents?
- The smarter the model gets, the more it overthinks everything
- Built a psychological portrait creation agent!
- Newbie to Agentic AI. What to look into next?
- Ten months running a weekly PDF product, and an honest list of which parts of the pipeline actually run themselves
- When an AI project starts going off track, what’s usually the biggest problem?
- Does an AI agent really need its own inbox? That's a very dangerous and architecturally wrong trend.
- What does a customer security team actually want to see before approving an AI agent?
- When should an AI agent question its own data?
- 51 AI API providers audited in 2026: 30 are still genuinely free
- [D] Self-Promotion Thread
- I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]
- Deepity: A C++ library showing Predictive Coding Networks can match Backprop (97.73% on MNIST in 60s) [P]
- I regret reviewing for AAAI [D]
- Where can I find legally usable datasets for advanced audio chord recognition? [D]
- Detailed explanation of how to create a text-to-image model from scratch. [R]
- I built an explainable bone-lesion screener for X-rays and ran it for £5 month [P]
- CABiNet (ICRA 2021) vs YOLO26-sem on UAVid: accuracy, compute, and GPU latency [P]
- MIR with AudioMuse-AI-SAE [P]
- Most open-source AI detectors can't hold a 0.5% false-positive rate [P]
- YOLO26-RGB: repurposing YOLO26's depth-trained backbone for image deraining [P]
- Latent Reasoning Landscape in 2026: Mapping BDH-CQ, HRM/TRM, Coconut [D]
- First A submission (AAMAS): how much theory is enough when your experiments went sideways? [D]
- What kinds of ML bottlenecks are a good fit for Triton? [Manning giveaway] [D]
- Are HMMs still used for unsupervised tasks? [D]
- We released TontaubeV1, a character-level TTS model for long-form generation [P]
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses [R]
- [D] Simple Questions Thread
- Can anyone explain how this works to me? Is it a scam? This person says they'll send me a computer and pay me $200 per week to keep it on 24/7
- Last quarter has been insane. Amazing times to be alive.
- What part of AI do you think we still fundamentally misunderstand?
- AI coding tools are saving me hours but I keep secondguessing whether I actually understand what I shipped
- What if tokens are not the giant labs' end game?
- Classics departments are disappearing and AI cannot even read their texts. The polytonic Greek problem nobody talks about.
- One unexpected way AI has genuinely changed my life: I repair things instead of replacing them
- Did anyone else notice Reactor’s new Orbis model? I tried turning it into an interactive game
- Study (n=504): heightened suspicion did not improve detection of AI-generated text, and fake-news accuracy fell 10.2 points under sustained exposure
- Used Story Prism’s New Agentic-Powered Building Tool to Connect 178 Sources in Minutes. Found a Disturbing Pattern in Epstein’s Intellectual Network...
- This university built an AI curriculum before ChatGPT. Now it wants to help other schools do the same
- Tools or Agents? Choosing Our AI Future - podcast with Anthony Aguirre
- Anthropic sued over alleged theft of 'tens of thousands' of songs | AI company faces multibillion dollar lawsuit over misuse of copyrighted songs to train Claude models
- I logged 240 generations of the same character over seven months and then measured how far her face actually moved
- AI chatbots tier list.
- Do you think AI agents need to become more accurate or more transparent about what they're doing?
- We have created an architecture built on top of Harness.
- ChatGPT's hard conversation-length limit is one of its most frustrating UX problems - even on Pro
- AI is the single most terrible invention in human history
- llm-gemini 0.34
- Claude's new system prompt really doesn't want to reproduce song lyrics
- Quoting Rick Brewster
- Claude Fable 5.1 made me a really nice animated pelican
- Codex bundles LibreOffice
- GeoJSON Map Viewer
- Quoting Tarn Adams
- datasette-mcp 0.2
- Python 3.15.0 candidate 2 is here!
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of Contributors
- [AINews] Fal’s H3 Max Live breaks the infinite videogen barrier
- Touchy
- Porte
- CleanShot 5.0 with Studio Mode
- Roadie
- Dyson CameraJet
- Basedash AI Sources
- Stitch AI by Dynamic Mockups
- RoundOS
- Doop
- Dynamic Edge
- GhostReply
- HONOR Robot Phone
- Cosmic Agent Plugins
- How Fast Can DeepSeek Run on 8GB VRAM? - HackerNoon
- DeepSeek and its peers are churning out chips at a frantic pace, has NVIDIA finally lost its shine? - 36 Kr
- Get 50 AI models and no recurring payments with a lifetime subscription to AskAnyModelAI Pro for $39.99 - PCWorld
- EXCLUSIVE: DeepSeek Was Just the Start. New ETF Targets China’s AI Tigers in a 'Convergence Trade' - Benzinga
- AI Helps Terror Groups Plan, Execute Attacks - Africa Defense Forum
- Commentary: DeepSeek out, SMIC in? TIME100 AI maps China's widening AI battleground - digitimes
- Beer Meets AI in China: This Beijing Bar Rewards Customers With Free Tokens - NDTV
- WeChat Pay expands AI AgentPay Card to DeepSeek Harness and OpenClaw - TechNode
- China AI Model Monitoring: Accelerated Upgrades, Surge in Usage, and Cloud Entering a New Upward Cycle - Moomoo
- Inside Z.ai’s turnaround after falling behind in enterprise AI - KrASIA
- China's Most Powerful Large Model: Who Is the Real No.1 Among All Claimants? - 36 Kr
- Tencent's Marvis Lets Users Plug In Kimi, Zhipu GLM and Other Third-Party Models - pandaily.com
- Southaven approves new SpaceXAI data center, fifth location in area - Action News 5
- Is Grok AI Safe? What to Know About Privacy & Safety - Private Internet Access
- Elon Musk Grok AI Predicts Ethereum Price by January 1, 2027 - TradingView
- ChatGPT, Claude, and Grok bypass EU restrictions on Russian state propaganda, RSF investigation finds - theins.press
- Telcos, Musk is coming for Voice - Sebastian Barros Newsletter
- Why analysts see up to 100% upside for SpaceX stock - TradingView
- Elon Musk publicly asks Grok for ‘vulgar’ roast of Billie Eilish - PinkNews
- Grok's Roadmap: What Musk's 'Only Gets Better' Means - BASENOR - Tesla Accessories
- Grok Bot Is Getting Cheaper: 4 Key Points on Token Optimization - BASENOR - Tesla Accessories
- The New Middle East After America – OpEd - Eurasia Review
- Silicon Valley Investor Says Grok Bot Sparks Agent Revolution as Monthly Fee Plunges 90%, Igniting Compute Race - finance.biggo.com
- Elon Musk Lauds Grok Bot: Why Silicon Valley Titans Hail It as the Next Revolutionary ChatGPT Moment - 36 Kr
- GitSpawn Flaws Let Malicious Repositories Execute Code in Claude Code, Codex, Cursor, and Grok - CyberSecurityNews
- 設定に1行足したら同じ一問が1.43倍安くなった。8分放置したら、その1行が11倍高くついた
- OCI Generative AIでGeminiもgpt-ossも呼べたので、7モデルの日本語力を比べてみた
- ナレッジ設計
- 「絶対に落ちないテスト」を1日に3つ作った。全部AIが自分で見つけた
- DeepSeek-V4-Flashを 145B に枝刈りして DGX Spark 1 台で動かしてみる
- LLMでドキュメントからFAQを自動生成する
- Claude Code のトークン節約でやっている 9 つのこと — 起動 7 万トークンの内訳と計測つき
- 「プロンプトを長くするのをやめろ」Fableのシステムカードから品質保証Agent Skillを作った
- 25億token学習したGPT-2 Smallに、さらに25億token学習させたら性能は上がるのか?
- Pythonで実装する量子回路の測定:確率振幅からビット列を抽出するプロセス
- Qiskitで実装する量子ビットの反転:Pauli-Xゲートの動作原理とビット反転のシミュレーション
- AIを褒めると性能は上がるのか?――文脈・メモリ・重み更新・RLHFを分けて考える
- 【試し読み】RAG構築入門 — LLMの「知識」という壁を突破する
- 立ち止まるきっかけは自分で作る──鵜呑み防止スキル discernment-nudge を使ってみた
- 完成版を学習させても精度は上がらない — 暗黙知は「作り方」で渡す
- LLMのツール呼び出しを型で閉じ込める: Valibot許可リスト検証とstrictObjectの罠
- ざっくりわかる AI Agent(5):重みを変えなくても、使うほどに強くなる——自己進化
- シミュレーターにAIを足すとき、計算だけはさせなかった
- AIの自動ルール更新で事故らない設計 ― “盲目的承認の罠”をHuman-in-the-loop×Gitで防ぐ
- Pangram 4 はいかにAI臭を見分けるのか
- numpyでAttentionを実装したら、√dで割るのは万能ではなかった
- Pokémon TCG AI Battle Challenge ポケカコンペ振り返りーメダルなし
- Instagramのリーチ予測モデル -「新規アカウント収集・モデル再学習」の自動化と監視を実装した
- YOLO26nアーキテクチャ徹底解析 step1 | 全体像
- 音声モデルを量産して分かった18のこと
- 「過去データは長いほど精度が上がる」は本当か?4時間入力が最も高精度だった | 第7回:AIで為替の未来予測は本当にできるのか?
- Instagramナノインフルエンサーのリーチ予測モデルを作った話
- ローカルAIモデル 2026 実践ガイド:メモリ別にわかる「動くモデル」の選び方
- GPT-2-likeからQwen2-likeへの実験:第1回 LayerNormをRMSNormに変える
- M4 Max 部署 Qwen3.8-27B:高性能推理与多人共享实践
- EVO-X2(Ryzen AI Max+ 395 / 128GB)でQwen3.8-Flash-Nextの高速化の続き。llama.cpp変更+MTPでデコード速度が約2倍
- Ingestシステムの構築と重み分散の最適化に向けて
- LLMに書かせた原稿の「差し込み漏れ」は、人のレビューでは止まらなかった
- 化合物の水溶解度を機械学習で予測してみる①
- 因果推論 Day 20/全30回 回帰不連続デザイン、閾値の際は擬似ランダム
- 深層学習をもう一段深掘りしてみた ~ニューラルネットワークの学習の仕組み・CNN/RNN/Transformerまで~
- 機械学習入門 第6回:PCAを「情報を保ったまま次元を減らす方法」として理解する
- 埋め込み行列はモデルから切り離せるか? 40,000 語の辞書を 2.4MB に圧縮して復元率 99.62% を実測した
- 【技術04】AIに「今どこ?」を聞かない。AI-OSのState Machine設計
- 【技術03】AIに会話を覚えさせない。Artifact中心でAIを動かす設計
- カブトムシの羽音――小さな命がLLMの内部体験をさせてくれた話
- なぜ「外側を守る」だけでは、これからの人工知能を守れないのか――人工知能の中に「免疫のような仕組み」を持たせる二重調整という考え方
- 「忖度なしで」というプロンプトは、ただのキャラ付けかもしれない。|LLM|ChatCPT|Claude|Gemini
- [Podcast] AIエージェントのトークン消費とプロンプトキャッシュについて話しました
- 🔊音声あり(日&英):【AIの守護神】たった0.5%のコストでLLMの危険をリアルタイム検知!SingProbeが未来を変える
- 「AIは結局賢いのか」という問いに答える前に読む本を書きました
- わかったつもりで書き、わからなくなって、また少しわかる
- Anthropicはなぜ本を裁断した?AI学習データの集め方と訴訟について
- 【生成AIニュース+】『Claude Fable 5.1 / Mythos 5.1』『Atlas』『DeepSeek-V4-Flash-Vision-Uncensored-GGUF』『video2dlssnr』『ComfyUI-H3VAE_TRT』『vh5tape VHS LoRA for MiniMax H3』『NKD Radial Menu』『MiniMax-H3 Fused Turbo INT8 ConvRot』『H3 Studio』他
- あなたのAIパートナーなら、この続きをどう書く?『親密さの文法』
- 166万円で「世界トップクラスのAI」を自宅に置ける時代になった
- AIとLLMの正体:初心者のための仕組み解説(第1回)
- 【「私」のバックアップが完了しました】脳の全てをAIに預けた未来。そこに残るドッペルゲンガーは「同じ私」なのか。
- 未経験でkaggleに出たら銅メダル穫れた
- 語る機械と、黙って感じる生命
- 【日記】llm Sim Life―小さな世界の観察記
- 同じ日本語文書がGPT-4oで3,557トークン、ローカルQwenで2,399トークン — 日本語ペナルティを4つのトークナイザで実測した
- ClaudeCode用のFable5.1の実践プロンプトガイド(公式情報参照)
- AWSをゲームで学べる「AWS Cloud Quest」に新バージョン「AWS Cloud Quest 2.0」登場! AIによるバーチャル顧客と対話し、要件を聞き出して正しくソリューションに落とし込め
- AWSとAzureが最大100Gbpsでの相互接続を開始。これでAWSはAzure、Google Cloud、Oracle Cloudとのマルチクラウドをサポート
- ボットの振る舞いを動的に学習して防御を改善し続ける、「Adaptive Intelligence」ボット検出エンジン、Cloudflareが発表
- ClaudeにSalesforceを統合した「Claudeforce」、AnthropicとSalesforceが発表。Claudeから営業データ分析や顧客対応を実現
- AIエージェントがRedshiftを操作してDWH構築や集計分析など可能に、Amazon RedshiftがAgent Toolkit for AWSと統合
- VS Code上で開発のセカンドオピニオンを別のAIエージェントから得られる「Rubber Duck」機能が実験的実装
- MIRU2026参加レポート