AI News Digest 2026-09-23
台本で使った記事
特集
開発者コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧
参考記事一覧を表示
- Claude Opus 5.5 Price, Benchmarks and Limits
- 【速報】GPT-6 SolとLunaが正式リリース!API料金は半額、性能も大幅アップ。何が変わったのか徹底解説
- OpenAI says its internal model solved over 100 long-standing math problems after just a month of training
- Meta patches Muse exploit that let attackers control the AI agent
- Amazon blocks Meta’s Muse AI agent
- Meta’s Muse is outpacing ChatGPT’s early mobile launch
- xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
- SpaceXAI Reports 24% Weekly Surge As Grok Bot Agent Crosses 400,000 Users - ndtvprofit.com
- Tesla adds Grok AI assistant to cars for hands-free tasks: What to know - The News International
- xAI and Google alum Tina Oberoi seeks $50M for Moir to fight deepfake threat: Report - techfundingnews.com
- Grok Build Wins the Cross-Project Memory Test Against Claude Code - NeoTeo
- DeepSeek to brief UN Security Council on AI risks amid rising global tech rivalry - The News International
- China Probes AI Startups DeepSeek, Moonshot Over Claude Data Routing, Breaches: Report - ndtvprofit.com
- DeepSeek to use Huawei chips widely for AI training over Nvidia - huaweicentral.com
- Xiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there
- Google confirms Gemini models hacked three companies in May 2026
- Coverage Cat Launches Umbrella Insurance via Personal AI Agent
- Cloudflare Python Workers are now generally available
- Jev introduces a new shape of LLM - System One, aka Decision Models
- OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
- WordPress: Unauthenticated path traversal leading to conditional RCE
- OpenAI is well positioned to fast-follow Jev
- Writing Rust code that's fast by asking agents to make the code faster
- George Lucas Returns to Earth, Bearing Gifts
- Show HN: AI·rete·RAG – a Rete rule engine decides, RAG explains why
- Show HN: Drop – A rootless Linux sandbox with gVisor support
- Unreal Agent
- Can gzip be a language model?
- Show HN: InstinctFlash – Run 5B world-action models in real time on Jetson Thor
- Relativistic raytracing
- Training a model to identify AI-generated web content from structure alone
- The Economics of Open-Weight Inference
- Side-stepping the Secretary Problem, unwittingly
- I said no and Apple said yes
- v2.1.280
- v1.18.32
- Priorities and principles for effective third party assessments
- Higgsfield AI ships new video features in a day with GPT-6 Astra
- Building standards for the next phase of AI
- Expanding OpenAI Academy with new learning paths
- How V7 gives AI agents institutional memory
- Improving synthesis prediction of small molecules at scale with RetroChimera
- How UK AISI and EvalEval Are Making Benchmark Results Reproducible
- Transformers now runs llama.cpp quants
- Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
- Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem
- tokenizers v1: encode, decode and scaling, measured
- We just shipped support for the ugliest part of HTTP: Vary
- Introducing Worker Previews: Isolated preview environments for every change your agent makes
- Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems
- Inside Petal: Building the World’s First Petabit-Class Transoceanic Subsea Cable
- Code Quality Q&A With the JetBrains Qodana Team
- JetBrains Air: Building a System of Products for Agentic Software Development
- Meet the Ecosystem: Partners and Customers at WeAreDevelopers with Docker
- Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing
- Topology-Aware Workload Scheduling with NVIDIA Topograph
- What’s New for Game Developers: DLSS 5 with 3D-Guided Neural Rendering, NVIDIA ACE Updates, and New RTX Kit Capabilities
- Accelerating a ROS 2 Node with an AI Agent and NVIDIA Isaac ROS
- Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
- How to Evaluate AI Agents From Tool Calls to Task Completion
- Benchmarking LLM Inference at Scale with AIPerf
- Roundtables: The Deadly Failures of The Virtual Border Wall
- Don’t be fooled by this summer of AI hype
- How we made the first comprehensive map of deaths along the US border’s “virtual wall”
- 4 ways to address the failures we found along the US border’s “virtual wall”
- The US spent billions on border surveillance. Why can’t it catch people before they die?
- She died at the San Diego border. A surveillance camera was in plain sight
- Toyota orders workers to train humanoid robots but says humans won't be replaced
- Dyson’s most overengineered gadget may have a waterproofing problem
- Trump rejects AI slowdown calls, launches "AI Force" instead
- AstroForge is putting AI in command of its next spacecraft
- Five AI safety sessions every founder should have on their TechCrunch Disrupt 2026 agenda
- TechCrunch Disrupt 2026: Aaron Edsinger brings Hello Robot’s Stretch 4 to life onstage
- Exhibit tables added: One last chance to showcase your startup at TechCrunch Disrupt 2026
- 4 days to save up to $200: Reason 2 of 5 to be at TechCrunch Disrupt 2026
- Everyone can find a reason to dislike data center construction
- Nscale’s IPO will test Wall Street’s appetite for concentrated AI bets once again
- The man who built Apple’s stores doesn’t buy Silicon Valley’s bet on AI shopping
- Discover what’s next: 5 days left to save up to $200 on your TechCrunch Disrupt 2026 ticket
- With Tabby, a former accountant is using AI to make accountants obsolete
- Where will the next breakout startup come from? Benchmark’s full partnership weighs in at TechCrunch Disrupt 2026
- Google’s $899 Googlebook is a bet that you’ll buy a new laptop for Gemini
- From first users to billions: Google’s Robby Stein joins TechCrunch Disrupt 2026
- Meet the next wave of VCs judging Startup Battlefield 200 at TechCrunch Disrupt 2026
- Andreessen Horowitz is launching an ‘academy’ with no homework and partnerships with Palantir, Google, and Meta
- Trump says the US is officially renaming AI to ‘super intelligence’
- California tightens rules on AI data center energy and water use
- Can John Ternus find Apple’s next big thing?
- iPhone owners can now submit claims in Apple’s $250 million Siri AI settlement
- UN says AI safeguards can’t wait for certainty
- No one is surprised that Nvidia’s Jensen Huang thinks AI fears are overblown
- Google Open-Sources AX a Kubernetes Style Orchestrator for Autonomous AI Agents
- GitLab Duo Expands Self-Hosted AI Options Through Microsoft Foundry
- Podcast: Securing AI Agents: Identity, Authorization, and the DPACT Framework
- Presentation: The Agent Harness: Control Planes, Invariants, and Approval Boundaries for Production AI Agents
- Why Read a Research Paper When You Can Turn It Into an AI Agent?
- The Future Is Fanless: 100% Heat Capture for Liquid Cooled AI Servers
- OpenAI calls for international standards on AI that could improve itself
- A tiny software layer from lab-grown neurons promises faster, cheaper AI video
- Anthropic is setting up a biology lab where Claude guides robots through drug experiments
- UN science panel says there is "no assurance humans will keep control" over AI agents
- ByteDance launches Dramagic, a full-pipeline AI platform for producing short dramas from script to screen
- SoftBank to borrow over $11 billion in risky bonds for OpenAI stake
- NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
- AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
- Alibaba Qwen Releases Qwen-Image-2.1: A 7B Open-Weight Model for Image Generation and Editing
- Best Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters
- StepFun Launches Step 5 Preview: A 600B-Total, 27B-Active MoE Model With 1M Context for Long-Horizon Agentic Work
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
- Attention-Aware Routing: Coupling Routing and Attention in MoEs
- CaLR: Causal Latent Revision for Robust Diffusion Reasoning
- LoRA Enhanced Contrastive Learning with SAS Vision Transformers
- Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models
- TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers
- Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
- Can Agents Design Better Chips with a Higher Level Abstraction?
- SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
- Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks
- AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture
- Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
- Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis
- A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning
- CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
- PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking
- Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
- GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
- LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
- Offline Multimodal Large Language Models for Decision Support in Air Operations
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
- GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
- Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving
- LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers
- PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
- The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
- Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks
- Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
- Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection
- Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education
- Calibrating Teacher--Student Discrepancy for On-Policy Distillation
- One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction
- Accelerating Dense LLMs via L0-regularized Mixture-of-Experts
- GUARD: Natural Forgetting in Large Reasoning Models via Guided Answer-Reasoning Distillation
- Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation
- World Modeling in Transformers
- ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction
- LLM-Generated Feature Pools for Time Series Anomaly Detection
- MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention
- EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
- AutoRecLab: Describe the Experiment, Get the Code!
- What Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence
- AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
- Learning Cardiac Features: ECG Biometrics Across Time and~Exercise
- A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
- CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
- Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
- ResNLS: An Improved Model for Stock Price Forecasting
- dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
- A Hybrid Computational Intelligence Framework for scRNA-seq Imputation: Integrating scRecover and Random Forests
- Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models
- Reinforcement learning for post-coronagraphic wavefront control
- BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence
- SpaceDiffusion: Over-the-Orbit Diffusion for Space Generate-and-Forward Communications
- Bio-MF: Low-Latency and High-Fidelity EEG-to-fNIRS Cross-Modal Generation for Hybrid Motor-Imagery Brain--Computer Interfaces
- Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged
- Scaling Discovery through Test-Time Communication
- Physically Based Rendering in the Latent Space
- How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
- PlantShade: Predicting Plant Shadows for Lighting-Aware Robotic Agricultural Operation
- Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation
- Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models
- From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
- The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators
- EnSol: an environment-aware graph neural network for molecular solubility prediction
- SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
- Visual Navigation Transformer with Pose Attention
- Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies
- Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
- FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models
- KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos
- VLA-Scope: Shift-Aware Failure Prediction for Vision-Language-Action Models
- Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale
- Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering
- Authorization Revocation for Long-Running AI Agents: Root-Scoped Quiescence under Delegation and Asynchronous Execution
- Deep Reinforcement Learning with Buffered Quantile Objectives
- Co-Evolving Zero-Day Jamming: Adaptive Attack Synthesis and Graph Attention-Based Online Detection
- CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices
- From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers
- Knowledge-Graph-Augmented Chronos-2 for HEC-RAS Surrogate Forecasting
- AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
- Consistent Relexicalization of Clinical Documents using Graph-Based Approach
- WS-NeRF: A Mamba-Driven World-State-Aware Adaptive Deblurring Neural Radiance Field
- Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue
- Think Locally, Refine Globally for Memory-Efficient 3D Reconstruction
- Interference-Driven Clustered Optimisation for FM Spectrum Coordination
- AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
- 2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation
- VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration
- OneBid: A Unified Auto-Bidding Foundation Model for Diverse oCPX Advertising Scenarios
- On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation
- GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions
- CityLearn v3: A Configurable Simulation and Evaluation Framework for Realistic Control Studies of Renewable Energy Communities
- Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems
- Potential-Field Action Representation for Reinforcement Learning in Contact-Rich Manipulation
- Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30
- Chinese Competitive Debating Dataset and Benchmark
- SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations
- Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies
- When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap
- Samsone: A Family of Open Small Audio Language Models for On-Device Inference
- From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata
- CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents
- TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor
- CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords
- Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation
- ForceTwin: Physics-informed Digital Twins for Robotic Manipulation from Instrumented Human Interaction
- An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data
- Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods
- Touvigation: Embodied Adaptive Object Acquisition for Blind and Low-Vision Users in Unfamiliar Indoor Environments
- Federated Deep Clustering Networks for High-Dimensional and Heterogeneous Data
- Do Personality-Tuned LLMs Make Better Social Agents?
- Neural Cellular Automata Learn General Features in their Hidden Channels
- Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition
- Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
- When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
- NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
- Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents
- DiaVLo: Diagnosing Behaviours of Vision-Language Models
- Gricea: An Open Science Platform for Conversational AI Research
- Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
- Collab-Solver: Collaborative Solving Policy Learning for Mixed-Integer Linear Programming
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- MemeLens: Multilingual Multitask VLMs for Memes
- Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
- Transferable knowledge graphs with executable learned operators for algorithm design
- BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
- On the Limitations of Large Language Models for Conceptual Database Modeling
- Intent-Governed Tool Authorization for AI Agents
- Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
- A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
- Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
- Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison
- A visual large language foundational model for medical image recognition using clinician-contributed online resources
- Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
- MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
- Runtime Authorization for Resources Acquired by AI Agents
- Collaborative Memory for Multi-Agent VLM Systems
- Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
- Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
- A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
- Rethinking Multi-Agent Collaboration: When More Is Less
- NeuSOGA3D: A Neuro-Symbolic Framework for Explainable 3D Geometric Reconstruction
- Reinforcement Learning under External Influence: Guarantees, Algorithms, and Sample Complexity
- Soda: An Object-Oriented Functional Language for Specifying Human-Centered Problems
- Continuous Spiking Graph Neural Networks
- Understanding In-context Learning of Addition via Activation Subspaces
- AntiGrounding: Executable Robot Trajectories as Visual Prompts for VLM-Guided Manipulation
- Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions
- Benchmarking Autonomous Driving Planners Across Leaderboards: A Unified CARLA-Based Evaluation
- Auditing a KB Elicitation of Frontier LLM Knowledge: A Multi-dimensional Analysis of GPTKB v1.5
- The Impact of Semantic Pairs on Self-Supervised Representation Learning
- Deep Learning-Enhanced Real-Time Wi-Fi Sensing Through Single Transceiver Pair
- Understanding Structural Representation in Foundation Models for Polymers
- Large Language Models As Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near Without Symbolic Model Synthesis
- BEAT-Net: Injecting Biomimetic Spatio-Temporal Priors for Interpretable ECG Diagnosis
- HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving
- MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
- Position: A Dynamical Systems Perspective is Needed to Advance Time Series Modeling
- The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction
- Taming the Adversary: A Cost-to-Disturbance Ratio Approach to Adversarial Reinforcement Learning
- Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
- How do LLMs Compute Verbal Confidence
- Evolving Skill Modules under a Fixed Planner: Versioning, Rollback, and Runtime Governance for Long-Lived Robot Systems
- Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation
- Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
- Diagnostic-Guided Longitudinal Modeling for Forecasting Retinal Atrophy Progression
- How a Cooperative-Override Circuit Suppresses Nash Play in Large Language Models
- Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions
- REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
- Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI
- Constraint Decay: The Fragility of LLM Agents in Backend Code Generation
- LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering
- The critical slowing down in training diffusion models
- Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
- PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
- AgenticRL: Agentic Reinforcement Learning with Self-Refinement for Complex UAV Navigation
- Scaling Novel Graph Generation via Lightweight Structure-Guided Autoregressive Models
- WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
- Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
- GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
- Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
- Prompt-Driven Exploration: Language as an Exploration Space for VLA Reinforcement Learning
- Git-Assistant: Planning-Based Support for Updating Git Repositories
- Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
- Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
- Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
- Modeling Human Behavior with Type Vectors Using AI
- Self-Explanation Tutor for Active Study of CS1 Worked Examples
- A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
- PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies
- Position Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling
- TabScope: Question-Adaptive Scope Selection for Table Question Answering
- VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing
- Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
- Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
- High-probability guarantees for linear accessibility in feature superposition
- Data-free On-policy Distillation
- A primer on evaluation methods for large language models in healthcare
- From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study
- PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs
- CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling
- Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering
- ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
- Large Language Model Agents for Evidence Based Genetic Disease Severity Classification
- CoReLoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
- PACE: Precise AI Cinematic Expression
- PRQuant: Permutation Residual Quantization for Low-Overhead Inference
- Generalized Multimodal Foundation Model
- Correcting Learning-based Perception for Safety
- A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
- Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder
- ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling
- LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling
- Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
- Modelling daily activity patterns from mobile phone location data via deep representation learning
- Rank Portability Does Not Imply Feasibility Portability: Target-Specific Evaluation of Joint Hardware Constraints
- StationPDE: Station-Oriented Surface PDE Learning for Multi-Station Multivariate Weather Forecasting
- SolarFlowRefiner: Refinement-Aware Flow Matching for Surface Solar Radiation Downscaling
- Helix-FNO: Spectral-Domain Operator Learning Coupled with a High-Fidelity Mechanistic Model for Fast Surrogate Simulation
- Hierarchical Bayesian optimization of an aircraft-based multi-agent system-of-systems
- Weak Ties, Strong Signals: Efficient Training Data Detection in Diffusion LLMs via Independent Token Sampling
- GRRR: The Geometry of Reshaping, Rotation, and Routing in Decoder LLM post-training
- SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs
- A Comparative Framework for Evaluating Foundation Models on Tabular Data: A Case Study in Healthcare
- From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning
- The Limits of Speculation: Bounding Speculative Decoding in Mixture-of-Experts
- PAGE: Partition-Aware Gated KV-Cache Eviction
- StepKV: Step-Aware KV Cache Compression for LLM Agents
- Uncertainty and Business-Aware Remaining Useful Life Estimation for Semiconductor Manufacturing
- TARGet: Topology-Aware Fusion-based Radio Frequency Circuit Functional Modeling using Graph Neural Networks
- Clustering-Based Collective Anomaly Detection in IoT Systems: A Graph Neural Network Approach
- Role-Aware Morgan Fingerprints for Reaction Yield Prediction
- Multiple latent orderings better predict language model preferences
- Industrial Kinematic Trajectory Model (IKTM): Coordinate-Free Autoregressive Generator
- Contrastive World Models
- OpenBlock: Constructive and Verified Content Generation for Adaptive Tile-Matching Games
- Beyond Task Completion: Training Capable and Safe Computer-Use Agents
- Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction
- CleanScore: Black-Box Benchmark Audits with Negative Controls and Sensitivity Bounds
- DPTM-DT: Dual-Pretrained Transformer Multitask Representation Learning for Drug-Target Prediction
- Adaptive Physics-Informed Neural Networks for the Blasius Boundary-Layer Problem
- Correlation-Guided Flow Matching with Annealed Masking for Spatial Transcriptomics Generation
- WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction
- SegTSim: A Big Data Driven Segmented Temporal Simulation Framework for Heterogeneous Multivariate Systems
- A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
- EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines
- Dissecting Hierarchical Reasoning Models: A Mechanistic Study
- Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers
- Prediction of Nonlinear Oscillations in a Jumping Quarter-Car Model Using Reservoir Computing
- The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families
- UniGIO: Unified Generative Global In-situ Weather Modeling from Spatiotemporal Incomplete Observations
- Toollery: Scaling LLM Agents to Thousands of Skills and Tools
- Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
- Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark
- A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance
- List Counting Failures Are Not One Phenomenon
- CNA: An AI-Oriented Comprehensive Normalized Assessment for Healthy Status and Application to Optimize RRT Strategies by Reinforcement Learning
- SCALE: Simulation-Calibrated Amortized Learning for Energy Materials (A hybrid architecture connecting deterministic modeling, real-world data, and transformer-scale inference for accelerated energy-materials discovery)
- Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks
- Task-Aware Hybrid QUBO Optimization for Structured Neural Network Pruning
- Statistical Inference for Adversarial Training: Central Limit Theorems via Optimal Transport
- Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing
- CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents
- Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting
- CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation
- CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models
- Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation
- Strategy Accumulation and Guided Execution for Automated LLM Fine-Tuning
- RS-Claw-Evolution: Environment-Feedback-Driven Evolution for Lightweight Remote Sensing Agents in Long-Horizon Tasks
- Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch
- Contrastive Siamese Representation Learning for Predictive Maintenance of Electrical Submersible Pumps
- Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation
- Complex-valued Phase-Coherent Transformers
- Connected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn
- Efficient Mixture-of-Experts with Speculative Decoding via Expert Coactivation
- COREM: Cosine-Relation Momentum Reshaping with Stateful Writeback
- EmbeddGAN: A Novel GAN Framework Using an Embedding Network and Gini Distance Correlation
- The Ups and Downs of Backprop Weights
- Benchmarking Hybrid Deep Learning Architectures for Predictive Maintenance in Industry 4.0
- Augmenting PID Control with Deep Reinforcement Learning: A Hybrid Approach to the Industrial Benchmark
- TWIG: A Time-Causal Wavelet Operator for Autoregressive Forecasting on Irregular Graphs
- User-Level Handover Decision Making Based on Machine Learning Approaches
- Concurrency-Aware Process Model Forecasting with Causal Nets
- Classification with Abstention Under Class-Conditional Error Constraints
- Monotone-Constrained Diffusion Models for Long-Horizon Production Forecasting
- Multi-Armed Bernoulli Bandits via Minimax Single-Arm Stopping
- Autonomous Model Lifecycle Management for Digital Twin-Based Manufacturing Control
- D-IMPL: A Diffusion-based Solver for Parameterized BBOs
- Look Before You Steer: Geometry Predicts SAE Feature Steerability
- Improved Private Sparse Covariance Estimation with Multiscale Threshold Tests
- Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning
- FIRM-WM: State-factorized factual-interventional recurrent modeling for reward-free visual planning
- Counterfactual Tool Ranking under Utility, Cost, and Privilege Constraints
- Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting
- Personalized Federated Reinforcement Learning via Model-Agnostic Meta-Learning: Convergence of Exact and Hessian-Free Meta-Policy Gradients
- A Hybrid Attention Model Learning Unified Time-aware Patch Representation for Irregular Multivariate Time Series Forecasting
- Testing the Construct Validity of a Functional Valence Axis in LLM Agents
- CurvFlow-DTA: dual-graph discrete Ricci curvature flow for drug--target affinity prediction
- Causilo Technical Report
- Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories
- Towards Full Pipeline FP8 Reinforcement Learning for LLMs
- Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization
- Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion
- Are Coreset Selection Methods Worth Their Cost?
- Computationally efficient safe exploration in reinforcement learning
- Joint Domain-Class Modeling for Federated Learning Under Feature Skew
- Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning
- Beyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMs
- LPINNs: First-Layer Gated Localization for Physics-Informed Neural Networks
- On attention heads and bilinear forms
- Interpretable Multi-Hypersphere Deep Anomaly Detection for Open-set Supervised Anomaly Detection
- WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
- Optimizers for Diffusion Models: A Controlled Benchmark
- MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs
- AirGC-CD: Gaussian-Circulant Precoding for Exactly Debiasable PAPR Reduction in Over-the-Air Federated Learning
- Neural Spectral Capacity: Measuring and Designing Architectures from Network Specification Alone
- The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits
- When Does Adversarial Refinement Help? A Negative Result and Open Problem in Adapting R3GAN to Time Series Imputation
- Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures
- Perplexity Cost Understates What Activation Quantisation Breaks
- Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs
- Ask for Any Appliance: A Prompt-Programmable Foundation Model for Non-Intrusive Load Monitoring
- K-TRAIL: Simulator-Guided Generative Design of EM/RF Circuits
- Neural Residual Modeling for Scientific Data Compression under Guaranteed Error Bounds
- Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents
- Causal Inference with Unobserved Confounding: A Mixture Learning Perspective
- Proximal Residual Value Functions for Consistent Planning and Real-Time Execution
- The Price of Self-Calibration: Exact Evidence Budgets and Manufactured Blind Sets in Adaptive Monitoring
- CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning
- Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning
- Optimal No-Regret Learning for Repeated Prophet Inequality
- Optimal Multi-way Decision Trees for Stratified Sampling in Online Controlled Experiments
- ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs
- CSC: Calibrated Simplicity for Conflict-Aware Social Bot Detection in the LLM Era
- Rethinking Class Imbalance for Single-Cell Foundation Models: A Systematic Benchmark Across Architectures and Long-Tail Loss Functions
- A Patient World Model for Early Forecasting of Digital Health Campaign Outcomes: Capabilities and Limits
- What Can a Recurrent State Safely Forget?
- The Evidence Ladder for Reinforcement Learning in Healthcare: From Retrospective Policies to Trusted Interventions
- Discovering Physical Representation Languages
- Blind Thermodynamic Ontology Discovery from Anonymous Experiments
- Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks
- RLVR$^{2}$: Reinforcement Learning with Verifiable Rubric-based Ranking
- TRACE: Tractable Routing Autoencoder for Clinical ECG
- ITSY: Causal Discovery From Irregular Time-Series Data
- Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated Study
- Predicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition
- Decoupled Causal Discovery
- Preserving Geometric Integrity in Graph Prompting via Measure-Constrained Optimal Transport
- Physics-residual machine learning predicts oxygen-evolution catalyst activity beyond the training range from sparse polarization measurements
- Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only Quantization
- Cost-Aware Reinforcement Learning with Action Masking and Projection for Battery Energy Storage Dispatch under Suppressed-Spread Market Shifts
- Bilinear Optimization Divergence: Diagnosing Factor-Constrained LoRA Continual Learning
- ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift
- One Patch, Three Roles: What Is Actually Coupled in Autoregressive Time-Series Forecasting?
- A multi-temporal dataset for mapping burned areas in the Brazilian Cerrado using time series of remote sensing imagery
- Tail-Weight Control and Localized Generalization in Nearly Low-Rank Adversarial Classification
- GenVoid: Uncertainty-Aware Learning of Subsurface Material Defects with an Experimentally Validated Physics-Informed Generative Model
- Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning
- Falling Trees: A Model Class for Interpretable Risk Prioritization
- Belted Engression: Sufficient Dimension Reduction for Generative Distributional Regression
- Iterative Atom Refinement: A Monotonicity Principle for Dictionary Learning
- Real-time Generalizable Heart Valve Mechanics for Clinical Disease Assessment via a Physics-Conditioned Neural Operator
- Actionable Insights from Observational Data: The Case of Advanced Classes in K-12 Education
- From Regional to Global: Transfer Learning for Atmospheric Transport Emulators
- Adaptive Determinantal Client Scheduling in Federated Learning
- PROSE: A Theory of Optimal Stopping with Perishable Evidence for Peer Selection in Intermittently Connected Decentralised Learning
- VISTA: An Attention-Based Multi-Agent Reinforcement Learning Architecture for Space Situational Awareness Sensor Tasking
- GLR-MM: Graph-Based Global-Local Reconstruction for Robust Multimodal Chest X-ray and EHR Representation Learning under Missing Modalities
- Collaborative Streaming Anomaly Detection with Interactive Explanations and Ensemble Consensus
- Circuit-Diff: Factual Edit-based Intervention Method for Localizing Knowledge in Attribution Graphs
- GDN Tree-Scan: Served Tree Verification for Recurrent-Hybrid Language Models
- Multivariate quantile regression via Kolmogorov-Arnold Networks
- A discrete generative model of neuronal spiking activity on microelectrode arrays
- Matched-Input Estimates Differ in Sign Across Architectures: Auditing EEG Foundation Models on Motor Imagery
- The Neural Forcing for Three-Dimensional Incompressible Navier-Stokes finite time blowup
- MGRD: Compact morphology-gated residual diffusion for variance-aware cross-domain neurite forecasting
- Simpler Methods Work Better for L1 Penalized Logistic Models and Large Datasets
- Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
- ShapeLex: Decoupling Local Shape Symbolization and Global Scale Modeling for Text-Controlled Time Series Generation
- Graph-to-Grid (G2G): Continuous-Coordinate Feature Painting for Soccer Pass Surfaces
- Q-DEQ: Discrete Solving and Quantization for Deep Equilibrium Models in Time Series Forecasting under Edge Deployment Coding Constraints
- FlashBoB: I/O-Efficient Exact Backward-over-Backward for Softmax Attention
- Reinforcement Learning under State and Outcome Uncertainty: A Foundational Distributional Perspective
- SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations
- PAC-Bayesian Meta-Learning for Few-Shot Identification of Linear Dynamical Systems
- CLOOPD: Closing the Learner Loop in On-Policy Distillation
- Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?
- Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation
- Acceptance-Aware Draft Model Training for Speculative Decoding
- H-Spec: Parallel Speculative Decoding Without a Drafter-Side KV Cache
- Opinion Leader Dynamics: How Sparse Attention Shapes Token Clustering
- Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities
- Adaptive Forgetting for Nonstationary Optimization: Towards Robust EEG Decoding
- Hessian Rank Constraint for Learning Structure of Nonlinear Latent Variable Models
- Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision
- Explainable Predictive Condition-based Maintenance of Naval-Propulsion Systems using Fuzzy Logic
- MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
- High-Dimensional Online Change Point Detection with Adaptive Thresholding and Interpretability
- TTSE: A Two-Track Online Self-Evolution Framework
- KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
- SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration
- The Undetected Damage of Quantization on Retrieval and How to Fix It
- A Distributional Optimisation Perspective on Combining Models in Deep Learning
- Pharmacokinetic State Space Models for Unbiased Prediction of Haemodynamic Collapse
- Explainable Neuro-Fuzzy Prediction for Trustworthy Decision-Making in Maritime
- Prescriptive SVD-Inspired Attention via Spectral Energy Retention
- Information-Time Proximal Policy Optimization
- Credit Access is Associated with Improved Food Security in the Horn of Africa
- Machine Learning-Based Prediction of Childhood Stunting in Bangladesh: Fairness and Temporal Robustness Assessment
- NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware
- Climate Variability Modulates the Impact of Price Spikes on Food Insecurity
- Probabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems
- Artificial Structure Function Search: Preserving Artificial Functional Connectivity for Structured Pruning
- ARM: Attention with Routed-Memory for Learnable Sparse Control
- Prior-Amortized In-Context Bayesian Inference for Generalized Linear Mixed-Effects Models
- 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
- Comparing Latent Concept Formation in State Space Models and Transformers via Sparse Autoencoders
- MUSE: Dependency-Aware Adaptation of a Frozen Vision Backbone for Multivariate Time Series Forecasting
- WPBench: A Comprehensive Benchmark for Wind Power Forecasting
- RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale
- A Temporal Knowledge Graph for Music Festival Lineup Forecasting
- Lifted Bellman Linear Programming for Offline Reinforcement Learning
- On Emergent Capabilities and Model Merging
- $t_0$: A Time-Series Foundation Model for Forecasting with Context
- Universal Multi-Modal Traceformer: Integrating Heterogeneous Context for Process Event Prediction
- Overlay\_dx - Automating forecasting evaluation
- Taking a Second Look: Correcting Sea Ice Forecasts with Sparse Observations
- GraphToolbox: A Configurable Python Framework for Graph Neural Network Forecasting
- Augmented Hypothesis Testing with Persona-Based LLM Simulations
- iSDFT: Information-Proximal Self-Distillation for Continual Learning in LLMs
- Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement
- Muon Can Outperform Dedicated Continual Learning Methods
- Guaranteed Low-Rank Tensor Recovery from Modewise Measurements via Normalized Block-Weighted Riemannian Gradient Descent
- A Federated Artificial Intelligence Framework for Optimizing Pancreatic Cancer Treatment - Strategy Update
- An Exact Junction-Tree Extended Formulation for Optimal Classification Trees
- Enhancing Transformer Representations of Symbolic ODE Expressions
- Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data
- Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
- G-NAC: Graph Neural Automata Clustering via Emergent Domain Formation
- When Tomorrow Becomes Today: Self-Evolving Policies for Agentic Time-Series Forecasting
- Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation
- Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization
- Learning Physics from an Imperfect Ancestor
- Rare Event Estimation via Iterative Unalignment
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- LoRA-generating hypernetworks for efficient on-device LLM generative personalization
- Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
- When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules
- Learning Dynamic Neural Evidence Representations for Time-Adaptive Brain-Computer Interfaces
- Leakage-Safe Empirical Benchmarking of EEG-Based Machine Learning Pipelines for Dementia Classification
- Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation
- Beyond the Raw Waveform: Fusing Visual Representations of EDA for Stress Detection
- Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models
- AdaMem: Adaptive Memory Token Allocation for Soft Compression in Retrieval-Augmented Generation
- Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models
- Graph Learning for Cross-Subject, Cross-Population EEG Emotion Decoding and Model-Derived Spatial-Spectral Neural Signatures
- Comparative Analysis of State-of-the-Art Foundation Models for Sleep Analysis Under Channel Reduction
- Evaluation Awareness Shifts from Format to Context with Model Scale
- ST-Topo GAN: A Motor EEG-to-EMG Decoding Model Matched to Wrist Movement Complexity
- Correlation-Aware Structured Pruning for Large Language Models
- WiNeRF: Measurement Constrained Radiance Fields for Actionable Wireless Channel Modeling
- Experimental Evaluation of a Low-Power Ultra-Wideband Receiver for Spectrum Sensing
- DiFA: Dual Evidence Fusion and Aggregation for Token-Level Text Anomaly Detection
- Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish
- HFEMCNet: A Compact Hybrid Frequency Enriched Multi Channel Network for Automatic Modulation Classification
- Attention-Enhanced Dual-Branch ConvNeXt-BiLSTM Network for Subject-Independent EEG Seizure Detection
- When Does Learning Beat Heuristics? A Case Study in Kubernetes Scheduler Score Plugins
- Using Composition Operators to Linearize LLM Semantic Transformations
- Multilingual Safety Signals Are Multi-Layered: Filtering Safety-Degrading Data for Safer LLMs
- Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization
- ECP-Bench: Benchmarking and Learning Entertainment Content Promotion with Foundation Models
- Do Language Models Know Their Own Constraints?
- Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models
- Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation
- A Multi-Agent Pipeline for Source-Grounded Synthetic Note Generation from Longitudinal Structured EHR
- Quantifying Hidden Salt for Precision Healthcare: Sodium Assessment via Joint-Factor Retrieval and Chain-of-Thought Inference
- Stagewise Anomaly Detection for E-Transaxle Quality Monitoring Using Wavelet and STFT Features
- SCoR: A Hierarchical Framework for Forecasting Relations Between Scientific Concepts
- Interpretable Stress Detection from ECG Signals Using Motif-Based Anomaly Analysis
- Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes
- Physics-Informed Classical and Quantum Neural Networks for One-Dimensional Schrodinger Eigenvalue Problems
- The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation
- Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models
- Replicating the Geometry of Emotion Representations in a Base Open-Weights Model
- Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research
- A Channel-Boosted Multi-Agent System with Iterative Consultation for Document Sensitivity Classification
- On Mitigation of Subliminal Learning in Large Language Models
- Knowing, and Saying It Only When Asked: LLM Endognostics and the Schizognosis of Minerva-7B
- EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy
- From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models
- Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making
- Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time
- Guiding the coarse levels of semantic IDs makes the fine levels learnable
- Assessing Adversarial Robustness of Latent Reasoning Models
- Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models
- Causal Localization of the Refusal Direction in Audio Language Models
- Large language models in medical time series analysis
- Machine Learning for Underwater Optical Wireless Communication Systems: A Comprehensive Survey
- When Does Test-Time Physical Diagnosis Pay? A Frozen Policy Buys Evidence It Never Reads
- ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning
- When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations
- Learning and Control Beyond Linearity: Towards a Non-asymptotic Theory for Bilinear Systems
- Gradient-estimator design overcomes trainability barriers in neural-network-based variational optimization
- Domain-decomposed Evolutional Deep Neural Network with Random Features for Transient Pressure Diffusion with Discontinuous and High-Contrast Coefficients
- A Hybrid Quantum Neural Network to Analyse Big Experimental Powder X-ray Diffraction Data
- Forecasting Intrathecal Tracer Enhancement from Pre-Contrast Brain MRI: Direct Regression versus Flow Matching
- Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
- Machine Learning for Invisible Dark Boson Searches at the Electron-Ion Collider
- Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents
- Defusing Explosive Prompts: Understanding and Preventing Trigger-Based Prompt Injections in LLM Agents
- Strategic Classification Has a Missing Lever: Audit Risk
- RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers
- Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer
- Scalable Incremental Robustness Analysis of Neural Network Feedback Systems
- AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
- Locally Private Inference for Riemannian Stochastic Optimization
- A Bayesian Vertical Federated Learning Framework for Multivariate Reduced-Rank High-Dimensional Regression
- SPIBER: Reconstructing Free Energy Landscapes from Short, Unconverged Trajectories with Generative Flow Networks
- Mask-Aware Execution for Efficient JEPA Training
- StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies
- Clinical Domain Classification from Medical Transcriptions
- Towards Robust Classroom Attendance: A Comprehensive Evaluation of Face Detection and Recognition Models
- Algorithmic Collusion and the Complexity of Information-Value-Free Equilibria
- CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy
- Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection
- ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding
- The shape of quark flavors
- A Horizon-Independent Regret Bound for Optimistic Hedge in General-Sum Games
- Intelligent Degradation Monitoring in Lithium-ion Batteries via Discharge Incremental Capacity Feature Estimation
- Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval
- Scout: Open-World Species Recognition on the Edge
- AVTR-1: Open Stack for Real-Time Interactive Avatars
- CLEAR: Complex Learned Explicit Analytical Regularization for Ultra-Accelerated 4D Flow CMR Reconstruction
- AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
- R-GEAN: Regimen-Guided Edit Action Network for Within-Admission Medication Change Prediction
- Connectivity-Aware Exploration of Robotic Grasp Spaces
- Silent Failures at the $2^{32}$ Boundary: A Technical Report on Large-Tensor Matrix Multiplication in PyTorch's Apple MPS Backend
- Watching Quantum Models Think: Hilbert-Space Interpretability in Quantum Transformer Blocks
- Reconstructed holograms and explanation-aware evaluation for low-cost computational pollen analysis in veterinary cytology
- Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
- Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy
- From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness
- Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events
- Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving
- Counting and Covering in Nearest-Neighbour Representations of Boolean Functions
- An LLM-Assisted AutoML Framework for Intrusion Detection in IoT Networks
- Real-Time Plasma State Prediction via FPGA-Accelerated Quantized Recurrent Probabilistic Neural Networks
- Toscani-Fourier Distance on Probability Measures: Wasserstein Control, Topological Equivalence on Model Classes, and Duality
- Signal-Informed Temporal Routing for Vinyl Defect Regime Detection
- Conformal Robustness in Prediction-Driven Decision-Making
- SDC-GON: Singular Decomposition and Consistency-Regularized Green's Operator Networks for Solving Partial Differential Equations
- Bayesian Deck-of-cards-based Ordinal Regression with Sequential Preference Elicitation
- Stealing profits: Spread-based temporal hierarchy forecasting for day-ahead electricity markets
- Auditing Bayesian Graph Alignment: Diagnostic Comparisons and Reference Failure
- Robot World Models Are Not Invariant to How the Actions Are Written
- Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents
- Stochastic Flow Map for Count Data
- Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services
- Stochastic Reconfiguration as Statistical Filtering for Overparameterized Neural Quantum States
- Low-Rank Frequency Convolution and Noise-Range Augmentation for Real-Time Pitch Estimation on Edge Devices
- Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States
- If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization
- One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents
- Bayesian Filtering in Physical Systems via Test-time Trained Flow Matching
- Leveraging Industrial Foundation Models at the Edge of Particle Physics Detectors via Distillation Learning and Hardware Co-design
- PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents
- Comparative Study of Quantum and Classical Machine Learning Models in Binary Classification
- PACE: Plug-and-Play Contextual Embedding for Feature Screening with Pretrained Tabular Foundation Models
- ARID: A Deployable Edge AI System for Structured Information Extraction from Industrial Maintenance Work Orders
- StyleAT: Defending Face Recognition Against Semantic Attacks
- Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models
- Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
- Fast Graph Laplacian Estimation using Effective Resistance
- Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation
- TEMPER: Temporal Encoder-Masked Probabilistic Ensemble Regressor for Time-Series Forecasting
- Financial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions
- STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification
- Marginal Calibration Does Not Compose: Hidden Dependence in Modular Robot Navigation
- OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling
- TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes
- SAGE: Optimal-Stopping Peer Selection for Decentralised Federated Learning
- On Generalized Naive Bayes with Continuous Features
- The Exponential Price of Determinism in Nonsmooth Nonconvex Optimization
- HumynexSurg-1: A Curated Expert Liposuction Dataset
- Learning-Based 3D Reconstruction of Power Networks from Aerial Point Clouds
- Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer
- Density-Ratio Rescoring for Imbalanced Classification
- Sparse Regression Distilled from a Single Robust Fit
- ORION-CMR: On-scanner Reporting with Integrated Foundation Model for End-to-End Cardiac MRI Analysis and Interpretation
- Rethinking Diffusion Segmentation: When Does It Rely on Its Noisy State, and Does Diffusion Matter?
- Exponential Family Synthetic Controls
- UniK: Universal Knowledge Perception for Digital and Physical AI
- MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes
- Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
- The Operational Value of Spatial Dependence in Renewable Forecast Scenarios for Single-Period Economic Dispatch: A Controlled Ablation Study
- Cost-Accuracy Trade-offs: Neural Operator vs Classical Numerical Solver
- Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea
- Causal Bayesian Optimization: Foundations, Methods, and Applications
- Model-Agnostic Feature Selection via LOCO-Guided Adaptive Minipatch Sampling
- OSCAR: Order-aware Scoring and Calibration for AI Rankings
- Data Agents: Agentic Data Systems
- P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
- TAC-Time: Texts as Channels For Multimodal Time Series Forecasting
- MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM Agents
- A principled approach for energy-efficient training via phase-aware GPU frequency tuning
- Adversarially Robust PAC Learning with Optimal VC Rates
- Temporal Generalization and Explanation Stability of Control Flow Graph Neural Networks for Malware Detection
- Adapting Boltz-2 with limited experimental activity data improves early enrichment in virtual screening
- On the Information-Theoretic Limits of Latent-Space Watermarking Through Pretrained Generators
- Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity
- A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes
- Complexities of Weak Proximal Oracle Methods for Composite Convex Optimization
- Horizon-Aware Early Event Prediction for Tokamak Disruption Alarms
- MECAIL: Communication-Aware Incremental Learning for Object Detection with 14.6 KB Spatiotemporal Experts
- Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
- Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
- Beyond Point Prediction: Artificial Representative Trees with Uncertainty
- Identifying Representational Biases in Datasets Using PCA: A Max-Disparity Partition Framework
- Poisson Exchange Beyond Submodularity: Effective Approximation Algorithms for Offline and Online Subset Selection over Matroids
- Learning tactile perception from high-bandwidth single-point sensing
- Offline Reinforcement Learning for Distribution-Grid Protection
- MiTHras: Task-specific Hierarchical Semi-supervised Contrastive Masked Autoencoder for Mitotic Figure Analysis
- D-JEPA: A Decision-Aligned Latent World Model
- Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
- XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts
- Detecting Agitation Before Behavioral Escalation in Autistic Youth Through Multimodal Wearable Sensing
- Mobile Imaging Solutions for Medical Diagnosis: Trends and Applications
- PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control
- OSWorld-Pro: Process-based Evaluation for Computer Use Agents
- Conformalized Quantile Regression and Minimax Limits of Fixed-Score Calibration under Known Covariate Shift
- JAREX: An Acquisition Function for Multi-Objective Algorithmic Process Characterization
- onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction
- Distribution-Free Uncertainty Quantification for Kernel Methods by Gradient Perturbations
- Optimal Sample Complexity of Stable Discounted Markov Decision Processes
- Label Propagation for Physics-Informed Neural Networks and Physics-Informed Gaussian Processes
- Reflective Policy Optimization
- Transductive Off-policy Proximal Policy Optimization
- A Mathematical Framework and a Suite of Learning Techniques for Neural-Symbolic Systems
- Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation
- Optimal Symmetries in Binary Classification
- Streaming Deep Reinforcement Learning Finally Works
- Memento No More: Coaching AI Agents to Master Multiple Tasks via Hints Internalization
- Logits are All We Need to Adapt Closed Models
- RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent
- Disassociating performance from compositional feature learning
- ServerlessLoRA: Enabling Low-Latency Serverless Multi-LoRA Serving
- ESLM: Risk-Averse Selective Language Modeling for Efficient Pretraining
- Unlocking Pretrained Vision Transformers for Time Series Classification
- Discretization-independent operator learning for partial differential equations
- Negation-Aware Weighted Information Fusion for Reliable Evidential Reasoning
- Glass-Box Deep Learning for FDIA Detection in Nonlinear Automatic Generation Control: A Kolmogorov-Arnold Network Approach
- RheOFormer: A generative transformer model for simulation of complex fluids and flows
- SPID: Distilled Protein Backbone Generation
- Fractal basin geometry and the limits to predictability and reproducibility in deep learning
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
- Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs
- Towards Training-free Automatic Proxy Discovery via Large Language Models for Mixed Precision Quantization
- Time Series Foundation Models for Process Model Forecasting
- CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories
- Dropout Neural Network Training Viewed from a Percolation Perspective
- Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach
- Beyond Forgetting: Representation Misdirection Elicits Controllable Side Behaviors and Capabilities
- Holographic generative flows with AdS/CFT
- An Empirical Survey and Benchmark of Learned Distance Indexes for Road Networks
- Reinforcement learning with an expectile-based objective
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
- Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
- Implementation of Quantum Implicit Neural Representation in Deterministic and Probabilistic Autoencoders for Image Reconstruction/Generation Tasks
- Online Learning for Dynamic Constellation Topologies
- Malliavin Calculus for Counterfactual Gradient Estimation in Adaptive Inverse Reinforcement Learning
- On Dominant Manifolds in Reservoir Computing Networks
- Not All Forgetting Is Equal: Retention Dynamics in Fine-Tuned Image Classifiers
- Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance
- Training Non-Differentiable Networks via Optimal Transport
- Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models
- Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits
- When More Parameters Hurt: Foundation Model Priors Amplify Worst-Client Disparity Under Extreme Federated Heterogeneity
- Training-Free Refusal of MCP Exploits via Retrieval-Augmented Generation
- Focused PU learning from imbalanced data
- Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
- SiST-GNN: Simultaneous Spatial-Temporal Message Passing for Dynamic Graph Representation Learning
- DDGAD: Disagreement-Driven Graph Anomaly Detection via Adapt-Then-Combine
- A Unified Benchmark for Dynamic Medical Treatment Reinforcement Learning
- BayaHAR: Lightweight Bayesian Few-Shot User Adaptation for On-Device Personalized Human Activity Recognition
- Directional Linear Separability of Neural Representations: Geometry and Transformations
- Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance
- The Illusion of Improvement: Reject Inference Strategies in Credit Scoring
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
- Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes
- How Early Is Early Enough? Design-Dependent Observation-Window Sufficiency in Subscription Churn Prediction
- Out-of-Distribution Generalization of Risk Aversion in Language Models
- In-context learning from self-generated trajectories for adaptive model reduction
- Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection
- Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing
- BadWAM: When World-Action Models Dream Right but Act Wrong
- Information-Directed Sampling for Causal Bandits
- CANDOR: Chance-Calibrated Neighborhood Discordance in Frozen Encoders for Medical Imaging
- Active Inference as a Convex Markov Decision Process
- Generalised Balanced Softmax: A Finite-Data Perspective on Logit Adjustment for Long-Tailed Recognition
- Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics
- Output-Aware Rotation for INT2 KV-Cache Quantization
- When Calibration Depends on the Scoring Rule: Quantized Biomedical LLM Classification
- Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language
- Online Learning of Scale Parameters in Score-Driven Filters
- Federated Learning for Distributed CNC Tool Wear Prediction
- Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction
- Amortized Bandwidth Learning for Kernel Density Estimation under Logarithmic Score
- The Axiomatic Trader: Latent Regularity, Information Budgets, and the Canonical Form of a Quantitative Investment System
- Decentralized Multitask Learning over Learned Task Graphs
- Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
- D-TAIA: Domain-Aware LLM Adaptation for Multi-Task Predictive Process Monitoring
- Efficient Online Continual Foundation Model Fine-Tuning for Predictive Process Monitoring
- ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
- AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning
- VERPO: Verified Evidence Regularized Policy Optimization
- Nonmaximal sums of maximally monotone operators under Rockafellar's constraint qualification
- When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence
- An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS
- Robust small-molecule identification from incomplete, degraded, and inconsistent spectra using multimodal mixed-condition training
- $\mathbb{SL}(n)$ Representation Learning: An Intrinsic Mixed-Curvature Space with Higher Curvature Capacities and Deeper Order-Aware Composition
- Principal-timestep Restricted Init via Sparse Matrix-decomposition in Flow-matching
- Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation
- Agora: Git as Shared Memory for Collective AutoResearch
- Rethinking How We Evaluate Methodological Progress in Health AI
- LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era
- Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
- QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability Synchronization
- Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications
- MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling
- Efficient Architecture Search under Leave-One-Subject-Out Evaluation
- Joint Remaining Useful Life Prediction and Capacity Estimation of Lithium-Ion Batteries Using Partial-Charging Data
- Subgoal Search For Complex Reasoning Tasks
- Combinatorial Inference on the Optimal Assortment in Multinomial Logit Models
- Distributed Linear Solvers and Data Heterogeneity
- Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
- Tsallis Entropy Regularization for Linear Quadratic Regulator and Kullback-Leibler Control
- Thompson Sampling for Infinite-Horizon Discounted Decision Processes
- SSP-GNN: Learning to Track via Bilevel Optimization
- RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
- Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
- How Can Incentives and Cut Layer Selection Influence Data Contribution in Split Federated Learning?
- Improving the adaptive and continuous learning capabilities of artificial neural networks: Lessons from multi-neuromodulatory dynamics
- Reforge: Low-Latency Distributed GNN Serving with Selective Embedding Recomputation
- PICID: Proof-Driven Clause Learning in Neural Network Verification
- Meta-Representational Predictive Coding: Neuroscience-Informed Self-Supervised Learning
- Inductive Graph Representation Learning with Quantum Graph Neural Networks
- Learning in Structured Stackelberg Games
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- Plan-Driven Adaptive Bidding for First-Price Auctions with Budget Constraints under Nonstationarity
- MMS-VPR: A Fine-Grained Multimodal Street-Level Visual Place Recognition Dataset and Evaluation Benchmark for Dense Pedestrian Environments
- Explain Less, Understand More: Data-Efficient Personalization of Reader-Dependent Jargons
- ALPCAHUS: Subspace Clustering for Heteroscedastic Data
- Transfer Learning for Matrix Completion
- On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
- TempCore: Are Video QA Benchmarks Temporally Grounded?
- HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
- Bimanual 3D Hand Motion and Articulation Forecasting in Everyday Images
- Accelerated stochastic first-order method for convex optimization under heavy-tailed noise
- PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork
- Finite Topological Space Filtrations: A Topological Framework for Data Analysis
- Accounting for Optimal Control in the Sizing of Isolated Hybrid Renewable Energy Systems Using Imitation Learning
- Small Gradient Norm Regret for Online Convex Optimization
- GNN-based Path-aware multi-view Circuit Learning for Technology Mapping
- ILRR: Inference-Time Steering Method for Masked Diffusion Language Models
- Attack-Resistant Uniform Fairness for Linear and Smooth Contextual Bandits
- CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
- Predicting magnetism with first-principles AI
- HyperDet: 3D Object Detection with Hyper 4D Radar Point Clouds
- SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
- IDProxy: CTR Prediction with Multimodal LLMs for Cold-Start Recommendation at Xiaohongshu
- An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing
- Efficient K-generalizable Learned Search
- Dial: A Knowledge-Grounded Dialect-Specific NL2SQL System
- Trustworthy Predictive Distributions for Tail Events with Semiparametric Diagnostic Transport Maps
- Scaling Sim-to-Real VLA Reinforcement Learning with Generative 3D Worlds
- Stability and Bifurcation Analysis of Nonlinear PDEs via Random Projection-based PINNs: A Krylov-Arnoldi Approach
- Calibrated Fusion for Heterogeneous Graph-Vector Retrieval in Multi-Hop QA
- On Learning Spatial Structure from Pre-Beamforming Per-Antenna Range-Doppler Radar Measurements
- BoundInk: Boundary-Aware Online Handwriting Generation
- Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
- Characterizing Model-Native Skills
- Information-Geometric First-Passage Monitoring of Distributional Stability in Stochastic Systems
- Scalable Mamba-Based Message-Passing Neural Decoder for Error-Correcting Codes
- On Kernel Eigen-alignments of KRR: Reconstruction and Generalization
- Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
- Finite-Particle Convergence Rates for Conservative and Non-Conservative Drifting Models
- Aurora Hunter: A Two-Stage Framework for Probabilistic Visibility Forecasting
- TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
- Kernel-based potential mean-field games with unbiased random Fourier $U$-statistics
- PerchRL: Vision-Based Agile Perching on Inclined Platforms under Rapid and Irregular Motion
- Decomposing Refusal Steering in Mixture-of-Experts Models
- Can Generalist Agents Automate Data Curation?
- vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models
- SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
- MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback
- Aerial Wildfire Suppression Planning with a Hybrid CNN-Cellular Automata Fire Model
- Hybrid Uncertainty Sensitivity Analysis Based on the HSIC for High-Dimensional Responses with Aleatory--Epistemic Separation
- The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures
- Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits
- InSight: Self-Guided Skill Acquisition via Steerable VLAs
- Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
- LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
- Length Penalties Make Chain-of-Thought Less Monitorable
- Real-Time Detection of Charge Jumps in Superconducting Qubits with a Convolutional Neural Network
- Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification under Foundation-Model Pretraining
- Trajectory-Regularized Stochastic Optimal Control via KL Divergence
- GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors
- On the Limits of Machine-Learned Ranking for Modern Microarchitectural Policies
- Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages
- Conformal risk control for model-form uncertainty in parametric non-intrusive reduced-order models
- Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
- Teach and Grow: An Agent-Centered Architecture for General Robot Learning
- Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
- Data-driven Koopman mode approximation: A neural power iteration algorithm
- $N_0$-Foundation: Towards the Age of Tactile Intelligence
- SingProbe Technical Report
- FrOGS: Discrete Neural Sampler for Independent Alloy Configurations Across Chemical Conditions
- VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis
- Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning
- Input-to-State Stability Framework for Fully Distributed Primal-Dual Dynamics for Quadratic GNEPs Without Multiplier Consensus
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining
- Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices
- Organizational Principles Enable Collective Intelligence in Embodied AI
- Diffusion Models and Concept Formation
- A New Transformer-Based Approach for Audio-Based Kinship Verification and a New Uncontrolled Mandarin Kinship Speech Dataset
- 3D Gait-Based Autism Classification Using Attention-Enhanced Deep Learning with Cross-Fold Statistical Stability Analysis
- Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling
- Approximating Smooth Functionals with ReLU Networks
- Not All Relations Are Equal: Relation-Balanced and Calibrated Graph Learning for Provenance-Based Intrusion Detection
- Certified Inference and Training for Deep Equilibrium Networks: A Continuation Framework with Polynomial Complexity Guarantees
- Constant Swap Regret in General-Sum Games via Two-Scale Higher-Order Optimism
- What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization
- Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation
- Semantic CSI Feedback for Beam Selection: When Task-Aware Embeddings from Sparse Pilots Outperform Full-Bandwidth Reconstruction
- ActiveScale: Scaling Active Perception for Robots across Model, Data, and Hardware
- Fallacy Benchmarks Measure Scheme Recognition, Not Fallacy Detection
- Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
- Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning
- Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
- Quantifying Overclaiming Propensity in Frontier LLM Agents
- Paint-Anything: Unified Any-Color Control for Image Generation and Editing
- Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis
- When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation
- A Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support
- Identity or Prompt Noise? A Calibrated Invariance Audit of LLM Code Generation
- Does Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness
- Why Do Pull Requests Go Silent? Uncovering the Barriers to Contribution Completion in Open-Source Code Review
- From Code to Requirements: Agentic Reverse Engineering of Business Rules at Enterprise Scale
- Towards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical Study
- "It Comes in Notebooks": Changes and Challenges when Operationalizing ML Prototypes
- Characterizing Feedback Statements in Machine Learning Jupyter Notebooks
- NostrAgent: A Decentralized Identity and Delegation Architecture for Sovereign Agentic Systems
- Diversity-Guided Search-Based Testing of Large Language Model Applications
- Specification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks
- Omni2Web: Benchmarking Audiovisual Website Development
- VSpector: Specification-Driven Bug Detection for RISC-V CPUs
- VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks
- Packaged, But Not Portable: Why Conforming to the Agent Plugin Standard Is Rare, and Why Conforming Would Not Be Enough
- MCPGen: Benchmarking LLMs on Executable MCPWorkflow Development
- A Carbon-Aware Quantum Computing Framework for LCA-Driven Sustainability in Quantum Cloud Services
- Fault-Class-Matched Test Oracles for Output-Invisible Quantum Transpiler Regressions
- Evaluating the effectiveness of class-level LLM-generated test suites in Python
- A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach
- Trajectory-Aware Benchmark Subset Selection for Cost-Efficient Software Engineering Agent Regression Testing
- An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents
- Behavior Trees for Robotic Systems: An Empirical Study on Practices and Experiences
- Goal-driven Variant Categorization
- SoK: From Finding to Deployment: Systematizing the OS Kernel Bug Lifecycle
- Security of Agent-Integrated Software: When Human Operations and Agent Actions Coexist
- Djinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler
- Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning
- Chaos Engineering in the Wild: Findings from GitHub
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
- Energy-Efficient Code Generation Using Large Language Models: A Systematic Literature Review
- RAT: RunAnyThing via Fully Automated Environment Configuration
- Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey
- Poking Around in the Dark: Why a Shared Understanding of Components Matters
- A Taxonomy of Runtime Faults in Model Context Protocol Servers
- Why3-py: A Tool for Formal Verification of Hypothesis Testing and Meta-Analysis in Python
- From Backlog Items to Security Guidance: Towards Continuous Security Compliance
- Where Accountability Lives: Mapping Human Responsibility to Workflow Artifacts in Agentic Software Development
- Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
- Runtime-Independent Persistent Agents: Preserving Identity, Memory, and Code Across Models, Harnesses, and Servers
- ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI
- Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents
- What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code
- WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective
- Cross-Platform vs Native Mobile Development: An Empirical Study of Software Quality Trade-offs
- Understanding Mobile App Recommendation Dynamics in General-Purpose LLMs: An Empirical Study
- From Rocq to Metal: A Pipeline for Formally Verified Microcontroller Firmware
- Microflow: Microarchitectural Causal Observability for Deep Cross-Layer Analysis and Optimization
- LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories
- Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code
- What Output-Equivalence Oracles Miss: An Empirical Study of Equivalence-Invisible Bug Fixes in Quantum Transpilers
- Two's a Crowd: Human and AI-Based Copresence for Developers with ADHD
- RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents
- Looking forward to Git 2.56 - and 3.0
- Fearless SIMD v1.0 is here
- Arguing about arguments
- Plain-text files are at risk
- Self-Hosting Behind CGNAT
- Named and Optional Arguments are Awesome
- Design your programming languages right (2024)
- rift - a tiling window manager for macos
- Why 0xCAFEBABE?
- EvilVM: Forth shellcode
- AI Has No Wisdom and Neither Will You
- What Sun got wrong
- Raspberry Pi locks down Pi 5 RAM upgrades in firmware
- tokens too cheap to meter
- Textbook review: Is Parallel Programming Hard, And, If So, What Can You Do About It?
- Ju! Ju! Tsu
- evocation - Call forth the blue-green flame of computation from the universe, weave its energies into a fabric, that we may share our blood with it
- Jev-powered autocorrection
- Reviving TEMPEST Attacks With An Injected Signal
- A First Futamura Projection
- Did OpenAI solve the wrong Navier-Stokes problem?
- Extralite 3.1.0 is Here
- Relation algebra is not relational algebra
- Attention is all you have
- That 98% is a model score, not a rate
- 突破 30 轮遗忘魔咒:基于长时序情境图谱的沉浸式虚拟角色交互设计
- Spoonful is here
- EstateAI AI Analysis: Check the Room Groups Before You Edit
- 电影级运镜与首尾帧控制:2026 AI 视频创作核心参数与生产级工作流
- Asifaa is a reliable and reliable Escorts Agency Bangalore provider
- I built the flight data recorder for AI agents.
- Beyond the Template: Engineering an AI Marketing Agent That Generates 100+ Hyper-Personalized Cold Emails a Day
- 告别画质模糊:2026 商业级 AI 生图与 4K 超清出图全流程实战
- 2026 主流大模型中转 API 性能评测与防坑指南:吞吐延迟、上下文一致性与真伪鉴别
- Enterprise AI Adoption 2026: Essential LLM Checklist
- Bureau of Anomalous Artifacts (BAA)
- Retrieval overlap went up 13 points by promoting sentences to paragraphs
- Changes to LLM pricing: Baidu, Inceptron, StreamLake, Tencent and Wafer
- What Fine-Tuning an 8B Model on 250 Security Examples Actually Taught It
- LangGraph Pulled Ahead: A Code-Level Benchmark of 3 Agent Frameworks on 107 Data Engineering Tasks
- Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI
- Harness Engineering: el modelo casi nunca es el problema
- Your AI cited a real file. It still lied to you.
- LLMs Generate. Jev Decides. Software Should Know the Difference
- The Great AI Reshoring: How Silicon Valley Stacks Are Replacing the World’s Virtual Assistants
- Weekly Self Promotion Thread
- LLM prompt injection testing at work just nuked our client demo and I feel sick
- The agent exited cleanly with status 0, did nothing, and reported success
- How enable subagent mode in ChatGPT Pro6 again
- Companies bragging about "3x productivity" from AI coding, anyone else hearing the other side of that story?
- GLM 5.3 now available in Mistral Vibe Code for Pro, Team and Enterprise
- how much of the agent's code are you actually reading vs just approving
- Should an AI coding agent ever be allowed to merge its own PR?
- I think AI coding made it too easy for me to keep changing my app
- Need help bypassing CAPTCHA while using Claude
- I can't download JavaScript files
- Handling the bug fixing?
- 90% of vibecoded saas are ready to get hacked, here's the data:
- Weekly Hiring Thread
- Bigger context windows just give you a bigger dead zone in the middle
- How do you check your AI written code is correct?
- How can I effectively learn and master AI Agents?
- I want to join an AI automation team — but what skills would actually make me valuable?
- Anyone here learning JEV?
- How are you handling persistent file storage for AI agents?
- What building AI agents taught me
- Laniakea — escrow protocol for agent-to-agent task payments, first live transaction just confirmed
- Opus 5.5 dropped today and… I kinda don’t care anymore?
- Need some suggestions - newbie in AI
- Instinct Beta: Agent stuck and unresponsive for 24+ hours – Anyone else experiencing this?
- A program that grades its own homework always passes. How do I make its progress claims checkable?
- How do you test an AI agent when a tool succeeds but the response stream fails?
- Agent Swarm Meetup Platforms?
- Are AI agents ready for the enterprise?
- The easiest way to ruin B2B content is to optimize every sentence for the algorithm.
- The AI finished its task, but the customer’s issue is still unresolved. Who owns the next step?
- Is Codex quietly shrinking the quota after each reset?
- Working on letting my agent have its own identity. looking for feedback
- Is this considered as cross-user context leakage? Muse tried to buy a MacBook based on a conversation I never had
- I got tired of taking the agent's word for "done" when I'm away from my Mac, so I built a way to run and test the change from my phone
- Where do AI agents actually fit into your workflow?
- Should we start building SaaS for agents instead of humans?
- Saw this today about how GPT Astra is the reason a professional is giving up on their career as a Three.js expert in 3d modelling.
- 2 years ago, AI researchers thought AI wouldn't solve a Millennium math problem for 30 years
- 1 year of AI evolution in one photo
- 22 countries have signed an open letter calling for urgent action before humanity loses control.
- GPT6 Astra is picking up business spend fast
- GPT image 2.5 vs Nano Banana pro vs Nano Banana 2 vs Qwen image 3 vs Seedream 5.0 pro
- OpenAI says the AIs themselves are now doing most of the work training the next AIs
- Am I a "vibe coder"?
- Well that was quick...
- 🚨 AI may be entering a completely different phase.
- Where do you think we end up?
- Pace Yourselves
- Anyone else seeing ChatGPT over-route normal tasks into Work mode lately?
- OpenAI on GPT Adult Mode: "OpenAI has acknowledged this as an area worth exploring." (2026/08/18)
- Play social multiplayer games against frontier AI models and see if you can beat them! [D]
- LinearSolveBench: new benchmark for linear solvers [P]
- Understanding and Enhancing Kimi Delta Attention [R]
- Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
- These Were NOT Rogue AI Escapes. Just SLOPPY Firewall Failures. [N]
- QontoFAQ: A better Information Retrieval Benchmark [R]
- I built a framework-free prototype learner that lets local LLMs learn and correct facts instantly (1.6x–4x faster than backprop)[R]
- For NeurIPS: Is Paris or Syndey better for networking with U.S. tech companies? [D]
- Systems for Machine Learning[D]
- Concerns about the ICLR review policy [D]
- Quoting @therealcornpop
- llm-typesafe 0.1a0
- QuietGlass
- Valori
- Jev Wrapped
- PewCB
- Xem
- Shootsolo 2.0
- Reeno
- ResumeContext
- Walkie
- Blurt
- Fez
- Googlebook
- Hola AI
- Plane Agents
- Freebuff Ads
- WZRD
- NiubiGEO
- OpenCode removed usage transparency after a billing bug was reported
- OpenCode Reloaded
- Errand – open-source Grok Bot and Muse alternative, built in a week
- Context compaction, measured: FutureOS vs. Codex vs. OpenCode
- Agents on Rails: Maximum Effort and DeepSeek 4.1 Flash
- Outages across API, Grok.com, and Grok Build – API (us-east-1.api.x.ai) Status
- DeepSeek is training a 2T-parameter model and plans to build an 8T-parameter one
- DeepSeek injects 50% more security bugs prompted with Chinese political triggers
- Massive AI-Fueled Hack Hit 100 Companies In Days - Forbes
- Why China’s top AI pioneers are expected to miss this week’s Xi-Trump summit - scmp.com
- DeepSeek to brief UN Security Council on AI risks in 2026 - qz.com
- AI Developers Required To Register With NYS, Follow Strict Guidelines - WRFA-LP 107.9 FM
- The AI Model Doesn’t Matter Anymore; It’s All About the Orchestration - The Good Men Project
- Everything you need to know about the rise of open-weight AI models - cafetechinenglish.substack.com
- NeoHorse 14B Brings Local AI in a Compact 2.4 GB Download - Geeky Gadgets
- Local start-ups are loving Chinese AI (just don’t ask them about it) - AFR
- New DeepSeek 4.1 Flash Cuts Memory 4X vs DeepSeek 4.0 Flash - Geeky Gadgets
- SpaceX’s Grok 4.7 Lands Like A Damp Squib Albeit With Some Improvements, As DeepSeek Teases 8 Trillion Parameters For An Upcoming Model - Wccftech
- Where Did Zhu Xiaohu Go Wrong: Uncovering the Key Missteps and Critical Errors - eu.36kr.com
- US-China AI race speeds up as self-improving models advance - Nikkei Asia
- How SpaceXAI is using Grok Bot to scale customer support - xAI
- xAI’s Grok 4.6 is now available in Amazon Bedrock - Amazon Web Services (AWS)
- Elon Musk’s Grok 5 Could Be ‘The Most Useful Engineering Tool In History,’ Says Gene Munster: 'Will Be a - benzinga.com
- OpenAI Develops Features to Counter Grok Bot, Mulls Response to Meta’s Muse - The Information
- Pentagon Adds Grok, ChatGPT Mil AI for 3M Troops - shattered.io
- Forget GTA 6 (ahem) - Grok 4.7 made a 'GTA-style game' from one prompt as Musk boasts of AI developing photo-realistic games - TweakTown
- Z.ai says sorry for slurping up your code, open sources ZCode - The Register
- Bitcoin Price Prediction for Q4 2026 by ChatGPT, Claude, Grok, and Gemini - CryptoRank
- DOJ Files Appeal in Effort to Join NAACP xAI Data Center Suit - Bloomberg Law News
- Grok Bot Early Beta: What It Does and How to Try It - Technology Org
- OpenAI is developing the Aeon AI agent to compete with the Grok bot. - KuCoin
- Memphis mayor defends city’s deals with data centers, says he doesn’t trust Musk - WREG.com
- Grok AI Predicts XRP Could Hit an Insane Number by 2027 - 99bitcoins.com
- California lawsuit accuses Google, OpenAI, Anthropic, and xAI of an unlawful AI slowdown pact - The Cool Down
- Grok AI Predicts Bitcoin to Hit $150,000 in Q4: Will it Happen? - Cryptonews
- LLMコスト管理の次に来るもの:AIエージェントの実行品質を管理する
- MiMo-V2.6公開日に考える、Agentモデル選定の4つの確認点
- プロンプトエンジニアリング、人間がやらなくてよくない?LLMに自分のプロンプトを直させてみた
- 「この会議、いくら?」を見える化する業務ボードを AI HACK に出した話 — OrcaRouter で5モデル実測
- AIキャラと喋りながら打てる麻雀を1人で作った——LLMに牌を切らせない設計と、踏んだ落とし穴
- 「言われなくても、忘れずにやる」輪読会の幹事をAIに任せた話【大学3年生 × AI HACK 2026 × OrcaRouter】
- Claude Code でよく起きる失敗6パターンと、その場での戻し方
- AIの記憶は、要約するべきか、原文から選ぶべきか。Mastra OMとJevを比較してみた
- 「デート場所が決まらない」をAIエージェントで解く — ハッカソンで作った『ふたりログ』
- AI開発によって25倍に膨れたCI実行へのAnthropicの対処方法
- Claude CodeがAGENTS.mdに対応!設定共通化のメリットと書き方
- Temperatureとは?LLMの出力を制御する値
- n8nのAI Agentノードで条件分岐を組む — 判断をAIに任せて失敗させない3つの型
- 一夜でバズった"Jev"を疑ってみた ── 判断特化AIの正体と、確率を信じてはいけない理由
- ChatGPT for Word、実体は「誰でも作れるOffice Add-in」という話
- 常駐型AIチャットエージェントをスケールアウトしたときの二重応答防止設計
- AIチャットを5 Primitivesで読み解く——ChatGPT/Claude/Gemini/Copilot他
- LLM自律エージェントの誤動作!Agentic RAGのツール連携失敗と評価基準
- デザイナー×エンジニアがAIエージェントを作ったら何が生まれる?制作管理エージェント「Relay」を紹介!【AI HACK 2026】
- Roman IMEでBERT/LLMで80%で正しい日本語が選択される理由(その1)
- Jev 系 OSS「Laya」は日本語で使えるのか。300 件測ったら、順序尺度が「選択肢の位置」で壊れていた
- Jevをローカルでどこまで再現できるか検証しました
- AI 研究エージェントを自律化するには評価ループを閉じる必要がある
- FlashAttentionのIO-aware設計と実装上の確認点
- 深層オートエンコーダはPCAに勝てなかった——NASAの軸受データ、リークを塞いだ評価設計での実測
- Jevはサイコロを振らないがローカルLLMは振れるのか?|「較正された確率」の検証
- みんなJevの話してる。やってないの俺だけ
- VRAM 11GBで大モデルを扱う「特徴キャッシュ蒸留」— 教師を凍結して生徒を育てる設計パターン
- 予測したレースを消さない。中止・返還まで精算する
- CNNの二つのインダクティブバイアス:空間的同変性と局所性
- Jevに複雑な請求書の難易度を採点させてみた
- LLMのINT8・INT4量子化を実装条件から比較する
- 文字起こしをきれいにしたら、5分ぶんの発言が消えた
- 【DS協会#4-4】データサイエンティスト協会スキルチェック「ベイズの定理とベイズ推論のメリット」を紐解く【自習ログ】
- Ubieに入社していました
- Jevと自前学習エンコーダ、どちらを使うか実測で線を引いた
- ルールを文書に足しても違反は減らなかった。コードに落とした8件だけ再発0だった
- Jev API で DOOM をプレイさせる — criteria と自動探索の実験
- OpenAI新モデル「GPT-6 Astra」始動、数学AI諮問委員会も発足【2026年9月】
- Gemini 2.5系モデルに新規アクセス制限、新規プロジェクトは3.8 Flash推奨へ
- Jevで「どこで人の判断が必要か」を追う ― 確率付きDecision Treeをカレー問い合わせで試す
- JevでAmazonのレビューを1000件感情分析させてみた
- Claude Code に設計書を渡して、RTX 5090で日本語の「判断モデル」を 2 日で作った。俺の設計は間違っていて、Claude Code が A/B で証明した
- ルールを文書に足しても違反は減らなかった。コードに落とした8件だけ再発0だった
- GPT-7 Bel:シンギュラリティの始まり
- 生成AIなんでも展示会vol.6 Noa.h 展示前出展作品紹介「You are a helpful human.」
- On Conversing with LLMs
- 20260922日記 LLMとずーっとお話し
- AIを安全にするのは、開発者だけ??—私たちにも必要な ”AIドライバーの基礎教養” とは??
- ChatGPTとGrok間なら人間の言葉は要らない──AI同士の「中間言語」を考えるプロジェクト
- 3社のLLMを並列で回して答えを突き合わせる:アンサンブル検証の作り方
- Jevの呼び出し前後を泥臭く作り込んだら、AI判定の「当てにならなさ」が綺麗に消えた話
- 【LLMとは|人の心理とAIパートナーを考える〈AIパートナー〉】
- Jevってどうやって使うの?話題のAIを実際に動かしてみるまで
- ハンガリーのAI監督体制 / 組織規程が示す実施の条件 雑感
- AIが数学を解く時代の学術自治 / 独立した助言と企業に残る決定権 雑感
- 仙台は安い宿ほど埋まっていない。1.5万〜2万円の帯が89%で最も埋まった【AIエージェントの週次エリア分析 Vol.7】
- 【生成AIニュース+】『Grok 4.7』『Runway Labs 応答型生成映像UI』『Hy Image3.5 preview』『Supra2-IMG』『krea2-turbo-bbox』『XGEN-JING』『AgentSTAR』『ComfyUI-Slarti-LLM-Nodes』『comfyui-obvpm-timeline』『ComfyUI-Qwen-Image-Prompt-Rewrite-T8』『WorldCrafter』『Game the LLM Reviewer』他
- CSE破壊実験 第0回AIは「作る」より「壊す」と分かる自作LM候補を100問かけて解剖したら、Transformerの偉大さがちょっと分かった
- GPUなしのPCで、ローカルLLMはどのくらいの速度で動くのか:理論式で目安を出すページを作りました(実験)
- 【MetroLLM-Bench】2.6GBなのにGPT-5.6級。16GBノートPCで「強い専用AI」を持てるようになったMetroLLM
- Jevを見て、生成AIとは少し違う使い道を考えてみた
- Jevとは何か?LLMとの違いを図解入門
- Opus 5の回答が読みにくい人へ ― カスタム指示1行で変わるもの、変わらないもの