AI News Digest 2026-09-10
特集
開発者コーナー
中堅コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧を表示
- GPT-6 Astra, looped transformers, and hidden reasoning
- Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits
- Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course
- SkillAdam: Stable and Efficient Skill Evolution for Agents
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- ExecCritic: Learn to Test, Test to Improve for Coding Agents
- Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI
- SE-GoS: Self-Evolving Graph-of-Skills for Skill Library at Scale
- MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging
- Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
- Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation
- API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
- MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents
- A-Evolve-Training: Autonomous Post-Training of a 30B Model
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- On the Navier–Stokes Millennium Prize Problem
- DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
- CVE-2026-82533: DeepSeek Harness Vulnerability Lets AI Agents Escape Their Own Sandbox - ox.security
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis
- ResidualAuth: What Authorization State Must Language Agents Preserve under Revocable Delegation?
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models
- Revoked but Still Authoritative: An Empirical Study of Revocation Enforcement in Agent-Memory Systems
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
- Towards Trustworthy Physical AI: From Theory to Practice Across Life Cycle
- $A^2E$ : An End-to-End Agent Auditing Engine
- A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
- AI Agents Push Humans Out of the Loop
- post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis
- Scanning the Harness: An Empirical Study of Supply-Chain Defects in AI Coding-Agent Configurations
- Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
- Suno replaces its AI models with a new one trained on licensed music as copyright suits pile up
- Presentation: Fixing the AI Infra Scale Problem by Stuffing 1M Sandboxes in a Single Server
- AWS is using Qualcomm for AI inference while Qualcomm uses AWS Bedrock to design the chips
- What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
- From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale
- Substrate-Portable Execution for Production LLM Workflows
- The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs
- When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems
- From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents
- Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
- Inference-Time Nash Alignment
- CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information
- Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment
- Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
- Agentic ML Exploration (A-MLE) for Ads Ranking
- Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation
- A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation
- CLAMP: Constrained Decoding for Vision-Language Embodied Planning
- PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
- Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
- When Agent Governance Helps
- Who Maintains Agent Skills? A Longitudinal Study of Human-Governed, AI-Assisted Skill Maintenance
- Detokenization Leaks: Reconstructing Local LLM Outputs From Cache Traces
- PCSDiff: Diffusion-Based Bias Correction and Super Resolution Toward Practical Operational Medium-Term Precipitation Forecast
- Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning
- Flow3D-OPD: Multi-Teacher On-Policy Distillation for 3D Geometry Generation with Flow-Matching Diffusion Transformer
- Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain
- AttnCompress: Dynamic Attention-Guided Trajectory Compression for Software Engineering Agents
- Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?
- The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?
- Learning to Configure Agentic AI Systems
- TimeWarp: Evaluating Web Agents by Revisiting the Past
- Hypergraph Enterprise Agentic Reasoner over Heterogeneous Business Systems
- Teaching agentic AI to generalize expert diagnostic reasoning in rare diseases
- AgentFairBench: Do LLM Agents Discriminate When They Act?
- FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness
- AI Snitches Get Glitches: Towards Evading Agentic Surveillance
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
- On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
- Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale
- Who Delegates to AI? Evidence from Agent Configurations in Github
- Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
- Logos: An Agent Harness on a Cross-Process Bus
- Localizing Emergent Failures in Agentic AI: Recovering Minimal Repair Families via Counterfactual Replay
- EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
- StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?
- Learning to Construct Practical Agentic Systems
- Grounded Skill Synthesis from Code at Scale for Agentic Intelligence
- SpecCoder: Specification-Aware Code Generation with Curriculum Dual-Task Reinforcement Learning
- Shortcutting the Fix: Identifying and Categorizing Agentic Exploits in Software Engineering Benchmarks
- "We Permit the Use of AI, but [...]": The Landscape of AI Policies in Popular Open Source Projects
- You can contribute if you... An Empirical Framework of AI Contribution Policies in OSS
- AWS、AIがフルスタックAWSアプリの基本コードを、セキュリティ、可観測性、インフラまで一気通貫で生成する「Nx Plugin for AWS 1.0」、オープンソースで公開
- AI に実装を任せるための仕様と検証ループ 〜ProxySQL に足りない機能を実装で埋めた〜 [DeNA インフラ SRE]
- Introducing ChatGPT Images 2.5
- How GPT-5.6 Sol helps run quantum computing experiments
- Automatic Key Exchange: faster, post-quantum secure origin handshakes for 45 billion daily connections (and counting)
- Get Gemini 3.8 Flash With 75% Off
- Hackers are stealing Claude tokens from subscribers
- Meta's Recipe for Building Agents as "Organizational Second Brains"
- AI Models Are Watermarking Text—Will You Notice?
- ASML locks in TSMC, Samsung, and Intel while Huawei races to break its grip
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill Abstraction
- Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
- AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
- AutoKD: Autonomous Knowledge Discovery
- Building Trustworthy Graph-Agentic RAG for Social Good: Architectures, Failure Propagation, and Assurance by Construction
- A Unified Policy Architecture (UPA): The Governance Kernel for Enterprise AI Operating Systems
- VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery
- A Hierarchical Consistency Framework for Auditing Retrieval-Augmented Generation Systems
- An Auditable Symbolic-RAG-Generative AI Architecture for Goal-Oriented Conversation Orchestration
- SkillAlign: Aligning Skill Interfaces for LLM-based Agents
- EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering
- BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents
- DART: A DAG-Based Reputation and Incentive Framework via Blockchain-Enabled Governance for Trustworthy LLM Multi-Agent Collaboration
- Structurally Close, Temporally Distant: Measuring Security Exposure in Long-Horizon LLM Agents
- Intent Drift at SME Scale: Deployment Practice, Not Model Capability, Determines Agentic Compliance
- SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness
- Event Interaction in Low-Rank Bottlenecks for Temporal Relation Extraction
- Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?
- Skynet: Workflow-Level Anomaly Detection for Agentic AI via Semantic and Structural Modeling
- WAPP: Safe Learning of Positive Security WAF Policies from Live Traffic
- Monadic Second-Order Logic in HOL: Deep and Shallow with Automated Faithfulness (Extended Preprint)
- CodeTD: Topology of Attention Detects Hallucinations in Code LLMs
- ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing
- What Eviction Destroys: A Restore-Counterfactual Audit of Forgetting in Agent Memory
- A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware
- Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method
- Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
- SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching
- Suan: Rectifying Direct Preference Safety Alignment in Large Language Models
- Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
- Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
- Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Developers
- Omni Interaction Agent Technical Report
- ReST-RL: Reinforcing LLM Reasoning through Unified Self-Training and Value-Guided Search
- AMA: Adaptive Memory via Multi-Agent Collaboration
- Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Outlier Detection
- TRACE: Trajectory Correction from Cross-layer Evidence for Hallucination Reduction
- Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version)
- Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-grounded Post-extraction Correction
- How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions
- Behavioral Grammar: Detecting Adaptive Malware via Tiny Language Model Priors and Second-Order Temporal Analysis
- ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
- Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
- Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
- Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
- LLMs Can Design Near-Optimal OR Algorithms
- Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation
- ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
- Benchmarking Language Models for Statistical Problem Formulation
- Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
- MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
- T-Mem: Memory That Anticipates, Not Archives
- Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study
- CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
- VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities
- NoLoCo: No-all-reduce Low Communication Training Method for Large Models
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
- Celty: SpMSpV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
- PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems
- Memory as Infrastructure: Reliability Engineering for Persistent Agent Memory in Months-Long LLM-Assisted Development
- The Impact of GenAI on the Future of Requirements Engineering
- Regret Dominates Surprise: Design-Time Requirements Engineering for Agentic-AI Safety
- SemVul: Semantic-Enhanced Graph Neural Networks for Code Property Graph-based Vulnerability Detection
- From Reading Code to Reading Spec: A Verified Layer for LLM-Driven Codebase Maintenance
- Authority Is Not a String: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents
- A Governance Methodology Layer for AI-Assisted Software Development: Defect Taxonomy, Controlled Ablation, and a Test of Process-Over-Capability
- SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control
- Python AI Agent Observability Without the SaaS Tax
- 生成AIによるデータ分析を「仕様」から始める:仕様駆動データ分析(SDA)の実践
- Xiao K Morning Brief | DeepSeek Significantly Cuts Pricing for Flash Series Models; Qualcomm and Amazon Plan to Jointly Build AI Data Centers - Moomoo
- DeepSeek's Huawei chip plan signals a shift to large-scale domestic AI deployments in China - globalsources.com
- Show HN: Self-hosted company OS, Claude Code and Codex agents in departments
- ChatGPT Sketch turns your bad drawings into detailed AI images
- Presentation: Platform Engineering in the Age of AI
- Hugging Face's new ML Intern lets anyone run machine learning experiments through a simple chat
- AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents
- EdgeMem: LLM-Free Agent Memory Construction and Retrieval via Evidence-Preserving Multi-Anchor Hypergraph
- Agents Trust Tools Too Much: Measuring Reliance on Unreliable Tools
- Inference-Time Graph Engineering for Multi-Agent LLM Workflows
- Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
- MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search
- Generator-Independent Runtime Assurance under Partial Observation
- DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
- SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use
- SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores
- Causal Attribution for Agentic Decisions: Estimators, Coupling, and a Traceability Specification
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool Actions
- RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving
- Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval
- Elastic Horizon: Discovering the Effective Interaction Frontier in Agentic Reinforcement Learning
- CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification
- Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression
- PRIMUS: Identity, Governance, and Verification for Multi-Agent Federations
- From Event Logs to Governed Action: A BlueSky Agenda for Agentic Process Mining
- Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans
- Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training
- WorldAgen: Unified State-Action Prediction with Test-Time World Model Training
- Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference
- FastE: Readout-Triggered Token Compression for LLM Embedding Inference
- SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs
- Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
- Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
- WolfSociety: Understanding Collective Risk from Harmful-Agent Scaling in Financial Agent Societies
- From Review to Authorization: Key-Isolated Threshold Signing for LLM Agents
- Versioned Transitive Dependency-Closure Binding and Operation-Time Effect Governance for Agent Skills: ClosureBound
- Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning
- All for 1-Bit: Towards Genuine 1-Bit Post-Training Quantization for LLMs
- SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
- Hardware Trojan Threats to Multi-Chiplet Photonic Neural Network Accelerators
- Characterizing Contention-Induced Reliability Collapse in KV-Cache Timing Side Channels for Multi-Tenant LLM Serving
- Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting
- No\=esis: Deterministic-First Retrieval with Two-Tier Context Hydration for Factuality-Critical Queries on Small Local Models
- A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
- AI for AI: Optimizing Additional Infrastructure Build-out to Power Artificial Intelligence Data Centers
- Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models
- MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models
- X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR
- Hyperparameter Scaling Laws Across MoE Sparsity
- Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling
- Large Language Models Transform Organic Synthesis From Reaction Prediction to Automation
- DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
- Localizing and Correcting Errors for LLM-based Planners
- Do Web Agents Investigate Before They Decide?
- When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
- To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models
- Subliminal Learning is a LoRA Artifact
- Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors
- Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
- UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
- TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
- FATS: A Prompt Injection Attack Utilizing Feign Security Agents with Deceptive Few-shots Learning
- Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
- PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
- Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
- Many-Tier Instruction Hierarchy in LLM Agents
- RAG over Thinking Traces Can Improve Reasoning Tasks
- PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
- LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
- AgentServeSim: Serving-System Simulation and Policy Search for LLM Agent Programs
- Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
- Limits of Reliability and Scaling in Language Models
- Where Is the Tradeoff in Using Third-Party API Routers for Agentic Software Development?
- A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
- Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression
- A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors
- Miles v0.1: Production-Level Post-Training
- HoneyRoute: Honeypot-Model Routing for Adversarial LLM Serving
- Squeeze10-LLM: Squeezing LLMs' Weights by 10 Times via a Staged Mixed-Precision Quantization Method
- Iterative GRPO: Batch-Online Multi-Turn RL via Single-Turn RLHF
- Unexplored flaws in multiple-choice VQA make benchmarking unreliable
- Backdoor Channels Hidden in Latent Space: Extending Cryptographic Undetectability to Modern Neural Networks
- VibeCheck: Assessing the Quality of LLM-Generated Unit Tests: A Multi-agent Empirical Study across Heterogeneous Repositories
- Beyond Lexical Metrics: Sentence-Embedding Detection of Reviewer Habituation in AI Code Review
- An Autonomy Aware Metamodel for Human AI Collaboration in Software Engineering
- One Is Not Enough: The Untold Story of Multiple Security Patches for One Vulnerability
- RepoNav: From Snippet Retrieval to File-Centered Repository Navigation for Code Agents
- Beyond Agent Harnesses: Cross-Substrate Authority for Multi-Agent Systems
- Understanding the (In)Security of Vibe-Coded Applications
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
- 10 Essential Claude Code Plugins to Upgrade Your AI Workflow
- An MCP Approval Server for AI Agents
- The Retrieval Pipeline Is Lying to You: How RAG Fails Before the LLM Sees Anything
- AWS、自然言語でデータ分析アプリを構築可能に、「Amazon Quick」に新機能
- Desert Ant Labs: local, fast models that run on device
- v1.18.30
- This AI entrepreneur is developing agents that can plan ahead for the unexpected
- Apple has a new way prove your iPhone photos aren’t AI slop
- Sequoia doubles down on Cymphony as AI agents create new enterprise security risks
- Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- AI Slop Is Changing How Engineers Review Code
- LWiAI Podcast #256 - Fable 5.1, Astra Tease, Gemini 3.8 Flash
- Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
- CriticGen: Generation-Aware Evaluation as Actionable Feedback
- When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
- SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
- Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment
- When and What to Teach: Budget-Aware Online Adaptation for Web Agents
- Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era
- EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent
- The Normalization of Deviance in AI Development
- DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
- Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment
- Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
- Agentic Pressure: The Endogenous Entropy of Reliable Autonomy
- Explaining AI Agents Through Execution Traces
- A Translational Note on AI Safety Evaluation
- SerenAI: State-transition system inspired by text-based world AI models
- Learning transferable human physiology from two million hours of sleep with SleepFM-2
- A visual large language foundational model for medical image recognition using clinician-oriented social media
- iBrain: A Unified Foundation Model Reading the Brain from Surface to Spikes
- Risk Is Not Review Value: Wrong-Answer Exposure Under Bounded Review Budgets
- AAS-RAIL: Improving Information Extraction for Asset Administration Shells through Retrieval-Augmented In-Context Learning
- AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
- What Does an LLM-Agent Leaderboard Rank Actually Compare?
- Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
- Beliefs and Behavior in Language Models
- A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate
- Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering
- Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models
- Personalizing LLM Agent Memory Using Biometrics
- GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data
- Deposon: An Auditable, Conservation-Guaranteed, Game-Theoretically Tested Scattering Layer over LLM Reasoning Paths
- An Agent Model Abstraction for Human-AI Teaming Cognitive Coupling
- Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding
- Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks
- Diamond Agent: Agentic Control of Federated HPC Resources as a Service
- It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center
- Robust Conformal Consensus: Multi-Agent LLM-as-a-Judge Interval Evaluation with Conformal Prediction
- SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
- Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning
- Mind the Phase: Effective Rank and Representation Health in Legged Locomotion
- Input-to-State Stability Framework for Fully Distributed Primal-Dual Dynamics for Quadratic GNEPs Without Multiplier Consensus
- Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Control System Anomaly Detection
- Distributed Lag Neural Additive Models
- Large-Scale User Behavior Analysis in Multimodal AI-Assisted Manual Task Execution
- Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs
- Kalman Delta Networks: Uncertainty-aware Associative Memory
- SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs
- LLMs for Social Network Modeling: From Network Generation to Dynamic Processes
- Tracing Stereotypes from Representation to Output in Multilingual LLMs
- Neptune: An AI model for Global Ocean Subseasonal Prediction
- Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems
- OntoKG-EQ: A provenance-grounded, competency-question-governed knowledge graph for auditable analyst querying
- Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation
- SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation
- Training-Free Task Vectors for LLM Behavioral Control
- Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
- Aligning Language Model Benchmarks with Pairwise Preferences
- NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
- A Progressive Training Strategy for Embodied Vision-Language Models to Mitigate Spatio-Temporal Hallucinations
- CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification
- Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
- ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
- Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
- Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI
- The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
- Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration
- KTO: Model Alignment as Prospect Theoretic Optimization
- HoarePrompt: Structural Reasoning About Program Correctness in Natural Language
- SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
- How to Backdoor Image Knowledge Distillation
- Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
- When Tools Hurt LLM Reasoning: State-Dependent Belief Revision under External Evidence
- GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initialization
- Boosting LLM Reasoning via Human-Inspired Reward Shaping
- DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
- Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
- Latent Preference Modeling for Multi-Session Personalized Tool Calling
- Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks
- PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
- Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
- Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
- HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
- Rollout-Level Advantage-Prioritized Experience Replay for GRPO
- QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
- Agentic Electronic Design Automation: A Handoff Perspective
- Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
- SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents
- Challenges and Recommendations for LLM-as-a-Judge in Multilingual Settings and for Low-Resource Languages
- Hierarchical Server Architecture for Agentic Science
- Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
- Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents
- On-policy Distillation with Verifiable Reward
- RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
- VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning
- AhaBench: Do Agents Learn from Prior Experience? A Benchmark for Long-Horizon Continual Learning
- MOLE: Detecting Insider Threats in AI Agents
- Broken on Arrival: Silently Defective LLM Artifacts in Public Model Registries and How to Catch Them
- Kronecker Factorization Improves Efficiency and Interpretability of Sparse Autoencoders
- Retrieval-augmented Decoding for Improving Truthfulness in Open-ended Generation
- ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages
- WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Small Pretrained Speech Recognition Transformers
- LGQ: Learnable Geometric Quantization for Image Tokenization
- Understanding Cross-Modal Contributions in Continual Vision-Language Models: A Theoretical Perspective
- Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining
- Interleaved Speech Language Models Latently Work In Text
- Secure Aggregation for Privacy-Preserving Federated Learning on Clinical EEG Data
- Pruned BPE: Post-training Visibility Pruning and Token Reallocation for Byte Pair Encoding
- LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
- Privacy Leakage in Federated Learning: Gradient-Based Client Identity Inference and Defenses for Inertial Sensing in Vehicular Edge Networks
- Correct Tests Are Not Enough: Measuring and Training Oracle Conversion in Specification-Based Test Generation
- KG-Commit: A Dynamic Knowledge Graph for Online Just-in-Time Software Defect Prediction
- Trust the Spec, Not the Code - A Specification-First, AI-Assisted Case Study in Online Banking
- Agent ATO: Visualizing Agent Interaction Timelines from Logs
- Reducing Hallucinations in LLM-Generated Code via Semantic Triangulation
- Towards Requirements Engineering for GenAI-Enabled Software: Bridging Responsibility Gaps through Human Oversight Requirements
- The best cross-AI memory tool: stop re-explaining yourself across ChatGPT, Claude, and Cursor
- I killed my fine-tune before I wrote a single line of training code
- Cut your Claude Code cost by 90% using the Spotify Method
- CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses
- Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
- Chinese AI firm DeepSeek taps underwriters for IPO: sources - South China Morning Post
- v2.1.266
- 1Password increases engineering productivity 21% with Codex
- IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
- Rust AI in Practice: Building LLM Applications With Rig
- Why this month's Microsoft patch release is a doozy
- Update to Google’s AI weather model improves forecast accuracy
- AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- Google Cloud races to catch up in the AI deployment wars with Accenture deal
- Microsoft has new AI privacy rules for schools
- Students who use AI generally score worse at school
- AI power users claim Anthropic duped them with subscriptions, and they’re taking it to court
- Adobe is trying to make its AI generators idiot-proof in Premiere
- More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
- Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization
- The End of AI Exponentiation: Fluttering Inside and Outside AI Bubble
- Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Review
- Reason Through the Latent! Making Latent Visual Reasoning Necessary
- NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
- When and Why LLM Causal Priors Help: Closed-Loop Prior Selection for Amortized Causal Inference
- EmoMed: An Emotionally-Aware Agent for Multimodal Medical Support with Real-Time Information Retrieval
- Agentic Algorithm Engineering: Improving Shared-Memory Exact Minimum Cuts
- World Models Under Asynchronous Sensor Observations
- The Internal Anatomy of Strategic Choice in Large Language Models
- A Tool-Augmented, GPT-4 Chatbot for Real-Time Repository Data Analysis
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
- xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
- CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows
- When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability
- Qiushi Engine on AstaBench E2E-Bench-Hard
- It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction
- RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems
- Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models
- SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement
- FACT: A Forensic Agent with Compiled Tool-Use Trajectories for AI-Generated Image Detection
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms
- Memory in Deep Time-Series Models
- ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs
- VERPO: Verified Evidence Regularized Policy Optimization
- Steering Geometry: Validating Human Value Geometry in LLM Steering Space
- ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language
- ECOKV: Geometry-Aware KV Cache Eviction via Complementary Diversity Metrics
- Frequency Estimation Based on SNR-adaptive Frequency Estimator Under Wide SNR Range
- MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation
- From LLM-Generated Specifications to Learned Quadruped Locomotion
- MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
- Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval
- Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval
- Quality Metrics for LLM-Generated Asset Administration Shells: A Perturbation-Based Evaluation Approach
- PLATOS: A Power and Latency-Aware Task-Oriented Scheduling Strategy for Healthcare IoT in Fog Computing
- Beyond Single-Negative Preference: Multi-Negative DPO for LLM-Centric Historical Entity Linking
- RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
- Latent-to-Latent Flow for Volumetric Stochastic Segmentation
- Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding
- Online Surrogate Repair: Decoupling High-Fidelity Feedback from Search Length in Closed-Loop Discovery
- How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
- Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
- Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions
- VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
- HyCO: A Hybrid Neural Solver for Combinatorial Optimization
- KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
- 3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints
- CS-CLIP: Compositional Scene Graph-guided CLIP for Robust Compositional Reasoning
- A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing
- RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation
- IPM-FM: A Foundation Model with Consensus Feature Selection for Industrial Process Monitoring
- AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation
- Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values
- Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-Constrained Settings
- From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
- SUN: Reaching for Novelty in Reinforcement Learning
- BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors
- Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning
- Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
- Adaptive Anisotropic Attention for Axis-Structured Signals
- Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports
- GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks
- It Is Not My Code Anymore
- Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
- TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
- Inferring the Unspoken: Aligning Embodied Agents with Implicit Preferences
- Neutralizing Popularity Bias in LLM-based Recommendation via Counterfactual Reasoning Guidelines
- The Emergence of Social Science of Large Language Models
- LOBERT: Generative AI Foundation Model for Limit Order Book Messages
- LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
- On the Context Sensitivity of LLM Moral Judgment
- Symbolic Informalization: Fluent, Productive, Multilingual
- Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design
- Walking on the DARKSIDE
- Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images
- EXAONE Finance 1.0: An Attention-free Time Series Foundation Model for Financial Time Series
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines
- Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
- Provable Pluralistic Alignment: Multi-Party RLHF under Offline Human Feedback
- Attribution in Scientific Literature: New Benchmark and Methods
- D-ADD: An Effective Plug-In for Defending Against Model Stealing
- Introducing HALC: A general pipeline for the systematic and reliable construction of prompts for automated coding with LLMs in the computational social sciences
- Positional Encoding via Token-Aware Phase Attention
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
- Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
- ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
- Agentic Inequality
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Backdoor Attacks in Federated Learning
- More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
- Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
- CARE: Confounder-Aware Aggregation for Reliable LLM Evaluation
- KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
- Adapting Technical-Service LLM Agents with Latent Logic Augmentation, Robust Noise Reduction, and Hybrid Reward Modeling
- PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset
- PAC-CF: Calibrating Irreversible Frontier Pruning in LLM-Guided Search
- AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
- Compliance vs. Sensibility: On the Reasoning Controllability in Large Language Models
- FragileFlow: Spectral Control of Correct-but-Fragile Predictions for Foundation Model Robustness
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- LitSeg: Narrative-Aware Document Segmentation for Literary RAG
- Models That Know How Evaluations Are Designed Score Safer
- The Little Book of Generative AI Foundations: An Intuitive Mathematical Primer
- Linear Separability of Activation Representations after Supervised Fine-Tuning on Incorrect Responses: A Study of Synthetic Dishonesty in Large Language Models
- Spike-Aware INT8 Execution for Spiking Language Models on Commodity CPUs
- When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer
- In-Context Multiple Instance Learning
- Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
- OdysSim: Building Foundation Models for Human Behavior Simulation
- Complexity and Scale in AI-Assisted Workflow Management: A Federated Learning Case Study
- Discovering Latent Groups for Robust Classification
- SEATauBench: Progressively Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
- TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
- Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
- Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
- Decision Making Needs Uncertainty Quantification [Lecture Notes]
- Probing Speaker Identity Sensitivity in Audio Deepfake Detectors
- WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
- SplitLite: Low-Rank Residual Compression for Split Learning
- When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study
- When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
- Co-Evolving Structured Knowledge and Reasoning in Language Models
- When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?
- ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
- From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
- EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
- Almost Free State Prediction Separation
- MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
- Robustness of LLM-Generated SystemVerilog Assertions to Semantics-Preserving RTL Transformations
- RAPTOR: Role-Aware Private Training for Mixture-of-Experts
- One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning
- Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems
- On-the-go Forgetting without Explicit Unlearning via ERASE
- Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment
- DataFlex-RL: An Evaluation Platform for RLVR Data Policies
- FMMO: Detecting the Divergence Between Local Attribution and Global Drift
- Hidden in Plain Sight: The Overlooked Significance of Canonical Elements for Extreme LLM Sparsity
- Data Efficient Sample Selection for In-Context Learning
- TrojanWorld: Backdooring World-Model Agents via Imagination Steering
- Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
- Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching
- MpSub: A Momentum $p$-Dimensional Subspace Trust-Region Method for Derivative-Free Fine-Tuning of Large Language Models
- MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference
- TV-Regulated OPD: Direction Matters in On-Policy Distillation
- From Narrative to Auditable Forecasts: A Structured Scaffold for Agentic Forecasting
- Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems
- LLM Layers Immediately Correct Each Other
- When Topology Betrays Privacy: Lattice-Based Reconstruction Attacks on Secure Aggregation in Decentralized Federated Learning
- CausalBN-Bench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
- SingLEM: Single-Channel Large EEG Model
- On Generalisation Error Bounds for Transformers
- Approaching the Harm of Gradient Attacks While Only Flipping Labels
- DDPM Score Matching and Distribution Learning
- Sustained Performance and Energy Accounting for Nonlinear Forecasting Across Classical and Simulated Quantum Models
- Decoder-Side Semantic Conditioning for Low-Bitrate Neural Speech Compression
- MGD: Moment Guided Diffusion for Maximum Entropy Generation
- The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
- Learn to Rank: Visual Attribution by Learning Importance Ranking
- An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility
- Physically Consistent Parameter Inference: Transparent Machine Learning Emulation in High Energy Physics and Cosmology
- A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex
- Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal
- Explainable Diabetic Retinopathy Classification Using Vision Foundation Models
- ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data Scarcity
- SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective
- When Stakeholder-centric Requirements Engineering is Not Enough: An Action Research Study on Legacy System Modernisation
- EnvPilot: Systematic Design and Evaluation of an Experience-Augmented Agent for Software Environment Setup
- Tool Retrievers Are Underestimated: Annotation Expansion Reveals True Capability
- Beyond Fixed Fault Models: Comparing LLM-Based and Rule-Based Fault Injection in OpenStack
- Measuring the Security of the Evolving Software Supply Chain: a Research Agenda
- A11yn: Aligning LLMs for Web Accessibility-Aware UI Generation
- OxyMake: A Content-Addressed Workflow Engine
- Do Code Language Models Follow Tests? Paired Interventions on Program Behavior
- The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness
- I Built SelfContext So I Could Stop Re-explaining Myself to AI
- NEW Ionic Framework MCP
- I ported Toyota's Lean quality system to Claude Code so the same agent mistakes stop coming back (MIT, free)
- I build my game with coding agents. The scarce resource is the decisions I still have to make.
- DeepSeek Used Distillation To Train Its R1 And V3 Models - Quantum Zeitgeist
- Disentangling Steering Vectors
- Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
- Think Wider: Mitigating Latent Rank Collapse in Implicit Chain-of-Thought Reasoning
- Anthropic scientist puts the odds of AI destroying humanity above ten percent this decade
- Muse by Meta
- AI Giants Work Hand-in-Hand With the Pentagon, Contracts Reveal - The Intercept
- Better AI code comment detector
- How we rebuilt Cloudflare Workers’ module registry for Node.js compatibility
- Superintelligence is coming. Should we let it?
- ControlAI’s Connor Leahy on why superintelligence is ‘not a weapon, it’s an adversary’
- China’s Regulators Take Aim at “AI Boyfriends”
- Patagonia has what AI data centers want, including no resistance so far
- ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models
- The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
- The convergent laboratory: when AI reasoning, autonomous experiments, high performance and quantum computing reshape chemistry
- CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
- Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
- LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
- CWF: A Collaborative Writing Framework for Personalized and Reliable Popular Science Writing
- From Concentration to Differentiation and Back: Routing Effective Rank in MoE Reasoning Cohorts
- SSP-DMGTimeNet: Physics-Constrained Learning for Spatiotemporal Trajectory Prediction of Vehicle Platoons
- EEG-Driven Decoding Framework for Passenger Hazard Perception in Highly Automated Vehicles
- Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
- Unraveling the Real Working Mechanism and Inherent Flaws of GAE: A Method for Interpreting Transformer Processes from an Economic Perspective
- Human-like moral judgments conceal divergent motive attributions in large language models
- Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision
- Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging
- Automated Design of Inventory Policy with Large Language Models: An Exploratory Study
- A Better Spur Should Start From Each Objective
- LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation
- The Surprising Effectiveness of Approximate Value Iteration in Self-Play
- A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes
- CrossModalQA: A Cross-modal and Multi-hop Benchmark for Multimodal Retrieval-augmented Generation
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails
- Newton Matching for Generative Modeling: A Unified Framework for Fine-Tuning and Sampling
- Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance
- Concord: A Video Relational Algebra for Cross-Modal Query Optimization
- Data Scout: Targeted Web Crawling for Domain-Specific Pretraining Corpora
- AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents
- What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction
- AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection
- STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient Large Vision-Language Models
- Solving versus Verifying: Catching Contradictions in Tax Reasoning Systems
- Beyond Cross-Lingual Transfer: Benchmarking Propagation Boundaries in Multilingual LLM Unlearning
- Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts
- FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon
- PAGR: Proof-Carrying Algebraic-Geometric Retrieval: A Quiver-, Provenance-, and Sheaf-Theoretic Framework for Grounded LLM Retrieval
- What the Window Does Not Contain: Auditing Provenance in a Document-Grounded Instability Benchmark
- SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs
- Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in EasyCrypt
- Collision Snapshot Guided Time-Reversed Safety-Critical Scenario Generation
- One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints
- A TTP by TTP Approach: Precise Malware Detection via Malicious TTP Recognition
- Certifying cooperation: a novel approach to cooperative multi-agent task generation
- Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models
- You Are What You Read: Misalignment via In-Context Persona Induction
- Constrained Online Learning with Noisy Constraint Values
- ARNAI: Artifact Removal Network based on Autoencoding and Inpainting for Robust Spinal Image Segmentation and Measurement
- Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
- AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing
- Mind the Approximation: Fisher-Weighted SVD Compression for ViTs
- Deep Learning for Biopsy-Free Subtyping of Basal Cell Carcinoma from Dermatoscopic Images
- TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning
- Parser-Free VLM Verification for Federated Weakly Supervised Video Anomaly Detection
- We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation
- Mapping the Emerging Social Science of Large Language Models
- ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making
- From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles
- Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
- Quantifying the Engagement Trap: Impact of Short-form Video Recommender Systems on Users with ADHD
- You Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs
- Foundation Models for Generalizable Semantic and Goal-Oriented Communication
- AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions
- TDDN: Text-aligned Diffused DINO Network for Puzzle Understanding
- Rethinking Sign Language Translation: The Impact of Signer Dependence on Model Evaluation
- DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning
- Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation
- CUNO: Curriculum and Preference Optimization for Stable Graph Unlearning under Mass Deletion
- CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations
- Segment Any Motion with Radar: Robust Multimodal Moving-Object Segmentation and Tracking
- Equivariance Breaks the Learning Rate
- SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation
- CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation
- CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling
- The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits
- DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
- NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
- Evaluating Steering Techniques using Human Similarity Judgments
- From the Fluency Fallacy to the Micro-to-Macro Validity Gap: Opportunities and Pitfalls of LLMs in Social Simulation
- What Does a System Modify When It Modifies Itself?
- Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
- A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads
- When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
- MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts
- Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
- MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
- PiMRef: Deducing Ever-evolving Spear-phishing Emails with Knowledge Base Invariants
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- BaNEL: Exploration Posteriors for Generative Modeling Using Only Negative Rewards
- CogniDir: Combating Cognitive Malicious Comments via Adaptive Distributional Learning for Robust Fake News Detection
- Deep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and Navigation
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data
- An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models
- TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
- TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
- AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
- CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
- AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
- AdaCultureSafe: Adaptive Cultural Safety Grounded by Cultural Knowledge in Large Language Models
- Attention-Weighted Value Projection for KV-Cache Compression
- AdaExplore: Failure-Driven Adaptation and Diversity-Preserving Search for Efficient Kernel Generation
- Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
- The Transformer as a Polar State Estimator
- Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
- Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models
- SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Encoder Spike Sparsity
- From Rashomon Theory to PRAXIS: Efficient Decision Tree Rashomon Sets
- Adaptive Calibration for Fair and Performant Facial Recognition
- HAARES Half-Split Residual Basis Routing for Deep Transformers
- What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents
- HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
- PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding
- CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA
- HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
- DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering
- CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
- Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
- IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views
- Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
- Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
- SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization
- Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps
- ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection
- Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
- Capsule Lens: Locating and Tracking Concept Geometry in Model Representations
- Online Learning with LLM Experts from Limited Feedback
- CALM: Class-wise Agreement and Label-gated Disagreement Modulation for Decentralized Federated Learning
- LoGIC: Budgeted Context Construction for Node-Level Graph In-Context Learning with Tabular Foundation Models
- Data Quality Rule Generation with LLMs
- FANS: Federated Adaptive Network Search Learning for Heterogeneous Devices
- Are Verifier Errors Independent Within a GRPO Group? Evidence from Qwen2.5 Rollouts
- Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs
- Continual Learning Mechanisms Compose for Long-Horizon Memorization
- The Oversight Gap: What LLM Safety Monitors Miss, and Why It Is Not Capability
- Dense Structural Compression of Transformers via Gauge-Correct Channel Removal
- On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing
- InfluenceField: A Differentiable Field with Interventionally Identifiable Causal Structure for Multimodal World Modeling
- Risk-Conditioned Fine-Tuning of Large Language Models
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
- Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation
- Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
- EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs
- Towards Bridging the Gap Between Offline and Iterative Alignment via Preference Distillation
- Content-Based Addressing for Long Context
- The OCUDU dApp Platform: An Open Runtime and E3 Interface for Real-Time AI-RAN
- Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
- CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
- TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
- Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics
- Evaluation of Contextual Understanding in Large Language Models
- A cautionary tale on the cost-effectiveness of collaborative AI in real-world medical applications
- Formal Bayesian Transfer Learning via the Total Risk Prior
- Differentially Private Model-X Knockoffs via Johnson-Lindenstrauss Transform
- Interpretable Network-assisted Random Forest+
- DAGLFNet: Deep Feature Attention Guided Global and Local Feature Fusion for Pseudo-Image Point Cloud Segmentation
- Residual-augmented flow matching operators for probabilistic partial differential equations
- Jointly Optimizing Debiased CTR and Uplift for Coupons Marketing: A Unified Causal Framework
- Learning Quantum Data Distribution via Chaotic Quantum Diffusion Model
- Hyperspectral Trajectory Image for Multi-Month Trajectory Anomaly Detection
- Hyperelastic constitutive model discovery with differentiable finite elements and structure-preserving neural networks
- Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction
- Atomistic Modeling of Chemical Disorder in Materials: Bridging Conventional Methods and AI-Assisted Approaches
- Amortized Neural Optimization for Pre-Layout Signal Integrity Design Space Exploration using Differentiable Surrogates
- Representation Costs in Data Science: Foundations and the Quasi-Banach Spaces of Deep Neural Networks
- Environment Parameter Gradient Theorem for Co-Design in Reinforcement Learning
- Learning-enabled Acceleration of Scenario-based Model Predictive Control
- Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images
- VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis
- Service Health Engineering for Distributed Systems
- Seeing Without Understanding: Large Language Model Evaluation of Mobile User Interface Quality, Failure Taxonomy, and Architectural Explanation
- Vectorizer: Vectorizing NumPy Programs with Shape-Guided Rewrite
- Detecting Multiple Semantic Concerns in Tangled Code Commits
- FIKA: Expanding Dependency Reachability with Executability Guarantees
- SynH-Rank: Quality-Aware Code Search via Diverse Data Synthesis and Hierarchical Ranking Training
- TianoForge: An Automated Bug Triage Approach for the TianoCore UEFI Firmware Development Community
- Defusing Logic Bombs in Symbolic Execution with LLM-Generated Ghost Code
- Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
- Enterprise AI Adoption 2026: Essential LLM Checklist
- How you frame a question changes what an LLM actually argues, not just its tone
- Deadline Work Does Not Belong on a Free Lane
- Not a traditional coder, but building a backend-guarded sales AI agent—how are you guys structuring the guardrails?"
- 1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install)
- Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- Infostealers Target Claude, Cursor, Codex and Other AI Agents to Steal Credentials and Sensitive Data - gbhackers.com
- Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- Towards Unified Multimodal Graph Foundation Model: A Bridge-Router-Adapter Based Approach
- Funding grants for new research into AI and teen development
- Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
- TypeScript 7 in WebStorm: Faster Coding Assistance for Angular and React, No Migration Required
- “This is the AI men actually use”: Meta ads pushed apps nudifying real teens
- Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- Viral AI assistant Instinct now has its own email address
- Amazon Prime Video’s new AI tech matches lips to dubbed audio
- Damage-Aware Bandit Pruning for Vision and Language Transformers
- RAPID: Reliability-Aware Pair Importance Distillation
- Planning and Scheduling Business Processes under Control-Flow Uncertainty
- Recovering Temporal and Geographic Signals from Language Model Embeddings
- Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration
- MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition
- MARBO: Relational Belief Grounding for LLM Agents in Social Deduction Games
- We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness
- Beyond Sparse Rewards: A New Benchmark and Structure-Aware Graph Alignment for Micro-Drama Understanding
- PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians
- Distance-Aware Attention and Wall-Distance Expert Routing for Transformer-Based 3D Flow Prediction
- Weakly supervised neural network: segmentation of complex structures in X-ray microCT
- DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning -- Extended Version
- From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction
- Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study
- Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture
- OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?
- Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation
- ProToMEx: Rapid, Interpretable Explanations via Structured Representations
- Situation Awareness for Intelligent Data Distribution in Connected Vehicles
- Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models
- PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement
- Closed-Loop Evaluation of Bird's-Eye-View Maps from Cross-View Transformers as Inputs to Behavior-Cloning Policies
- Where Does the Sound Go? Tracing Acoustic Information Loss in Audio-Conditioned LLMs
- Grounded and Faithful P&ID Reasoning: Constraining Vision-Language Models with Recovered Evidence Graphs
- DART: Distributional Adversarial Recurrent Training for Algorithm Learning
- Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning
- Image-Scale Robustness and Visual Recognition Performance: A Cross-Architecture Analysis
- From Splats to Silicon: Rethinking Computational Efficiency of 3DGS
- Linear Algebra Foundations of Efficient Attention: A Phase Reversal in Rank Collapse Under SVD Compression
- On BatchNorm Forward Modes in Value-Based Reinforcement Learning
- Tracking the Moving Frontier: Long-Short Term Advantage Estimator
- Comparative Study of Anatomical and Learned Features in AI Models for Structural Brain MRI
- The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists
- AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories
- LoGAN: Multilingual Font Localization with Generative Agents
- Temporal Heterogeneous Graph Transformer for Credit Card Fraud Detection
- FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity
- In-Place Instruction Following in Diffusion Language Models
- Protocol effects on feature-based hardware-Trojan detection across Trust-Hub families
- Parallelism Strategy Chaining for Fast Training Convergence
- Mathematical Programming in Machine Learning and Artificial Intelligence: A Unified Taxonomy of Models and Applications
- Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing
- Human mutation field reveals an equilibrium-like structure with irreversible circulation
- Topology Obstructs Pure Foundation Neural Quantum States
- Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?
- An emancipatory vision for designing (generative) AI for learner flourishing
- Climate-ModernBERT: Revisiting Corpus Composition for Domain-Adaptive Continued Pretraining
- The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]
- SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
- Sparse Data Augmentation for Optimization with Provable Guarantees
- WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos
- Exploring Bottom-Up Clustering for Creating Semantic IDs
- Leveraging contextual events on structure-aware next activity prediction
- TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection
- GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI
- Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors
- Optimal Experiments for Partial Causal Effect Identification
- Robust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A Review
- Evidential-Based Higher-Order Set Argumentation Framework
- Precision at Scale: Domain-Specific Datasets On-Demand
- Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
- Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation
- ChatBEV: Empowering Traffic Scene Understanding and Simulation via Vision-Language Model
- Clinician-Friendly Foundation Models for Ophthalmic Image Diagnostics without Fine-Tuning or Technical Barriers
- Evaluating the Scalability and Adversarial Generalization of GRPO-Trained NLI Models
- A Theoretical Analysis of Provable Compositional Generalization in Neural Networks: A Necessary and Sufficient Condition
- Knowing Your Uncertainty -- On the application of LLM in social sciences
- Hypersolid: Emergent Vision Representations via Short-Range Repulsion
- Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models
- Parity, Sensitivity, and Transformers
- BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR
- ASDA: Automated Skill Distillation and Adaptation for Financial Reasoning
- Can We Change the Stroke Size for Easier Diffusion?
- Building evidence-based knowledge bases from full-text literature for disease-specific biomedical reasoning
- CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
- LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback
- Gradient-Free Training of Spiking Neural Networks via Low-Rank Evolution Strategies
- Improving Federated Graph Recommendation with Semantic Guidance
- Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts
- Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
- Adaptive Densification for High-Fidelity and Efficient Sparse Gaussian Splatting in Arbitrary-Scale Super-Resolution
- Resonant Brane Splatting for Arbitrary-Scale Super-Resolution
- An AI-Based Decision-Support Pipeline for Day-Ahead Photovoltaic Forecasting
- Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability
- Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
- Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation
- Token-Oriented Semantic Communication with Pretrained Vision Transformers
- When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
- Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine Learning
- Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models
- CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
- Text-guided flow matching enables sample-efficient crystal structure generation
- StateSwap: Probing Support-Elimination Hidden States in Multiple-Choice Questions
- Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports
- How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI
- Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity
- Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning
- Machine Learning for Pre-Culture ESBL Risk Stratification to Guide Empiric Antibiotic Selection: A 12-Hospital Study of Enterobacteriaceae Cultures
- Rethinking One-Shot Federated Graph Learning: Training-Free Statistical Estimation
- EgoNeMo: Transferable Map of Pedestrian Dynamics via Egocentric LiDAR Scan
- When Retain Constraints Conflict: Mitigating Forget-Retain Interference in Tabular Data
- Exact Record Omission in Delta Attention: A Transport Criterion, Its Cost, and a Replay Certificate
- Fine-grained Distributed Backdoor Attacks in Federated Learning
- Latent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics
- Solving the Elastic Wave Equation with Physics-Informed Neural Networks: A Robust and Critical Assessment
- Proactive Context-Forecasted Safety Constraints for Nonstationary Reinforcement Learning
- Distillation as Probability Transport: Routed On-Policy Distillation
- Do Reasoning Representations Help Humans Evaluate LLM Outputs?
- PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games
- Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane
- Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache
- Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model
- Decomposing LLM-Judge Uncertainty to Target Expert Labels
- RoPE attention is an exact forward-pass gradient step with softmax intact
- Dynamic-Programming-Guided Hierarchical BPE and Empirical Analysis of Vocabulary Pruning
- LatentMD: Benchmarking Markdown Boundary Failures in LLM-Generated Text
- EigenLI: Spectral Approximations to Late Interaction
- To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
- Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
- Proper Dataset Valuation by Pointwise Mutual Information
- The Dynamics of Generalization in Deep Learning
- A Review of the Long Horizon Forecasting Problem in Time Series Analysis
- Leveraging Discrete Function Decomposability for Scientific Design
- Differentiable Causal Discovery for Singular Linear Models under Confounding
- PatchFormer: A Patch-Based Time Series Foundation Model with Hierarchical Masked Reconstruction and Cross-Domain Transfer Learning for Zero-Shot Multi-Horizon Forecasting
- Convergent Stochastic Training of Multi-Headed Attention and Understanding LoRA
- Can Revealed Preferences Clarify LLM Alignment and Steering?
- When Can Conformal Risk Control Certify LLM Outputs? Bounds, Impossibility, and Adaptation for Structured Generation
- The EM-algorithm and the Method of Moments in Softmax Mixture Models
- ControlTac: Scaling Tactile Data with Physically Controlled Tactile Image Generation
- Mind the Gap: Navigating Inference with Optimal Transport Maps
- Dynamical stability for dense patterns in attractor neural networks
- Learning Acrobatic Flight from Preferences
- Learning Multi-Index Models with Hyper-Kernel Ridge Regression
- Synthesizability Prediction of Crystalline Structures with Structure-Aware Feature Learning and Uncertainty Quantification
- Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer
- Amortized Inference for Correlated Discrete Choice Models via Equivariant Neural Networks
- Conformalized Super Learner
- Geometric Dictionary Learning of Dynamical Systems with Optimal Transport
- Pre-Warm: Initializing Convolutional Filters from First-Batch Patch Dictionaries
- Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data
- Bridging Ab Initio Symmetries and Global Nuclear Masses with Interpretable Neural Networks
- Any-Dimensional Learning by Sampling
- Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level
- Sequential operator learning under dependent data
- A convolutional framework for detecting event-driven dynamics in energy price series
- Replications, Revisions, and Reanalyses: Managing Empirical Evidence in Software Engineering
- FPScan: An Automated Constraint-Based Analyzer for Floating-Point Anomaly Detection
- RefVerifier: Semi-Automated Reference Claim Verification for Scientific Manuscripts
- A Large-Scale Dynamic Characterization of Flaky Tests in Quantum Software: The Qiskit Terra Case Study
- Email, decoupled from where and how: A cross-runtime, cross-provider email library for JS & TS
- A Design Space Exploration of Async/Await
- Goose + Perplexity 404 on Every Chat: the /v1 Trap and the Fix
- what's your approach to bus factor when the person who owns the code can't explain it either
- Deepseek Has Soft Retired Deepseek V4 Pro
- SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX
- DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision
- DeepSeek V4 Pro 0813 Scores 53, Prices Jump 3.6x [2026] - shattered.io
- Amazon Linuxが4年ぶりにメジャーバージョンアップ、「Amazon Linux 2027」パブリックプレビュー。SELinuxがデフォルトで強制モードに
- A Theoretical Framework for Masked Pretraining (MPT)
- One Step, One Lead: Mitigating Higher-Order Interference in Multi-Domain Reinforcement Learning via Cross-Step Control
- Feature Superposition in Neural Networks: From Theory to Practice
- Particle Dynamics of Flow Matching and Classifier-Free Guidance from a Stagewise Geometry Perspective
- NeuCME: Toward Dynamic Multimodal Continual Learning via Neural Combinatorics of Multiple Experts
- A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay
- Streaming Hierarchical Inference with Tabular Foundation Models
- Two-Scale Localized PCA-Net: Coarse-Global and Local-Residual Representations for Artifact-Reduced PDE Operator Learning
- Learning Metamaterial Eigenmodes with Wavelet-Encoded Fourier Neural Operators
- PhysSAE: Mechanistic Interpretability with Sparse Autoencoders
- Temporal-Causal Inference for Reinforcement Learning via Automata Learning
- I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models
- Prevalence calibration as shortcut mitigation
- $\alpha$-Graph: Attention-Infused Normalizing Flow Approach to Tractable Graph Modeling
- Sharp Structure-Agnostic Minimax Risk for Partial Linear Models
- GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra
- Geodesic-informed Generative Diffusion Model For Topology-preserved Image Video Generation
- DeepSeek's Next AI Battle Could Come With a $75 Billion Valuatio - GuruFocus
- I haven't lost a customer service fight in seven months
- OpenAI expands initiatives to support journalism from classrooms to newsrooms
- Man told ChatGPT he was feeling delusional. ChatGPT insisted he was Jesus.
- Apple CEO John Ternus says the best AI device is still the iPhone
- Compiling VGDL into Causal Models
- PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectories
- Distilling Vision-Language Models for On-Device Fire Understanding
- Exposing Weaknesses in Emotion Recognition in Conversations
- Generating Instance Generators in PDDL Planning
- Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts
- A Computational Implementation of a Goal-Directed Theory of Affect
- Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agriculture Platform in Hainan
- Simulating the Marginal Green Contribution of AI Modules in a Smart-Agriculture Platform: Evidence from Two Monte Carlo Experiments
- Unsound Search with Policy and Value Networks in Legends of Code and Magic
- RAFM-SER++: A Lightweight Multimodal Emotion Recognition Framework for Real-Time Behavioral Monitoring in Surveillance Systems
- Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End
- A radiographic world model for clinical reasoning and evidence generation
- TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs
- CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring
- Application of curiosity driven exploration methods for hardware interference identification
- When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic
- Emergent Goal-Directed Attention in Large Vision-Language Models
- Dual-Latent Memory Routing for Vision-Language Reasoning
- Adaptive Cost-Sensitive Machine Learning for Autonomous Robot Navigation Failure Prediction: When Not All Errors Are Equal
- GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims
- AVSplat: Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning
- The Role of Gradient Modification in Heavy-Tailed Nonconvex Stochastic Min-Max Optimization
- Explainable Deep Learning for Price-Trade Dynamics: From Black-Box Forecasts to Effective Parametric Models
- ExpertLens: Visualizing Embedding Spaces for Post-Hoc Explainability in MoE Enhanced Retrievers
- Decision-Aware Suffix Prediction and Reasoning of Business Processes
- VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
- Second-Order Smooth Planning with Optimal-Transport Bellman Smoothing
- Power Mean Estimation in Stochastic Continuous Monte Carlo Tree Search
- Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability for Small-Object Detection
- Mind the Gap: Exposing LLM Translation Blind Spots Using the AlphaMWE Multilingual Parallel Corpus
- SwiftExplorer: Training-free Diffusion Model Alignment with Swift Diversity Exploration
- A Trustworthy Watermarking Framework for LLM-Generated Food Safety Content
- DrugReason: Dynamic Multi-View Reasoning over Knowledge Graph and Language Evidence for Drug Repurposing
- Typed Federated Artifacts for the Agentic Web:Sharing Tool-Routing Knowledge Across Frozen,Heterogeneous LLM Agents
- XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
- Human-agent discovery of reconfigurable in-plane ferroelectric superdomain control
- SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
- AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis
- REFINE: Trajectory Representation Learning via Closed-Loop Transcription -- Extended Version
- LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances
- BlueprintAgent: Constraint-Triggered Targeted Revisits for Simulation-Ready Generation from Scanned Structural Blueprints
- Revisiting Thinning Methods for Kernel Learning Problems
- Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
- Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining
- Efficient Exploration Is Enough
- Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
- TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking
- IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA
- Synergistic Fusion of Topological Structure and Temporal Semantics of Mobility for Urban Region Embedding
- FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy
- Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
- ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR
- Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting
- From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation
- Instance-wise Linearization of Neural Network for Model Interpretation
- DUA-D2C: Dynamic Uncertainty Aware Method for Overfitting Remediation in Deep Learning
- Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering
- Q-Guided Stein Variational Model Predictive Control via RL-informed Policy Prior
- Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution
- Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
- SFO: Learning PDE Operators via Spectral Filtering
- Can One-Shot Test-Time Data Augmentation Help with Generalization?
- MAST: Mask-Guided Attention Control for Training-Free Regional-Multi Style Transfer
- Partner-aware Peptide-Protein Interaction Prediction and Target-conditioned Peptide Generation
- Deep Learning-Based Segmentation of Peritoneal Cancer Index Regions from CT Imaging
- Semantic Context-aware mOdality fUsion Transformer (SCOUT): A Context-Aware Multimodal Transformer for Concept-Grounded Pathology Report Generation
- HeteroGenManip: Generalizable Manipulation For Heterogeneous Object Interactions
- WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting
- EPC-3D-Diff: Equivariant Physics Consistent Conditional 3D Latent Diffusion for CBCT to CT Synthesis
- EmoTrack: Clinical-Semantic Modeling for Text-Based Depression Severity Estimation
- Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
- SymTRELLIS: Symmetry-Enforced Voxel Latents for 3D Generation
- TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation
- Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout
- ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
- Turing's First Imitation Game: Design Concepts and a Human-Approximates-Machine Reading
- DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology
- How to Verify Probabilistic Consistency of Predictive Models
- VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
- Generative Action-Chunk Sampling for Adaptive Stiffness Control in Physical Human-Robot Collaboration
- Assessing Alignment and Stability of Feature Importance Explanations via Weight of Evidence
- EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection
- Connecting Score Matching, Maximum Likelihood, and Expectation-Maximization in Mixed Linear Regression
- Generalizing HVAC Control With Domain Randomized Reinforcement Learning
- Beyond Arbitrary Geometry: Topology Generalization In neural PDE Operators
- Budgeted Task-Aware Acquisition of Dynamic Networks
- Learning to Price and Stock Under Contextual and Censored Demand
- Connectome-to-Function: Conditional Generative Latent Representations for Reservoir Computing
- Spectral Prioritized Sweeping in Nonstationary Reinforcement Learning
- SeaCausal-FL: Federated Fuzzy Causal Learning for Maritime IoT Fault Diagnosis and Counterfactual Reasoning
- How Does Parameter Pruning Reshape DNN Representations? An Interaction-Driven Exploration
- Not Just Oversmoothing: Detecting the Echo Chamber Effect in Graph Neural Networks
- Conditioned Initialization for Attention
- Revisiting Spectral Representations in Generative Diffusion Models
- Geometry-Aware Bayesian Parameter-Efficient Fine-Tuning on the Stiefel Manifold via Stein Variational Gradient Descent
- MLIP Detective: Active Failure Mode Discovery Beyond Benchmark Scores for Machine-Learning Interatomic Potentials
- Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families
- Chimaera: A Mixture-of-Graph-Experts Architecture for Cross-Task and Cross-Dataset Graph Learning
- The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation
- Learning Length-Extrapolatable Recurrent Models
- Benchmarking Storage Systems for Machine Learning Workloads Using NIO Bench
- SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
- ModularPhaseNet: Finite-Cyclic Phase Geometry for Computable Semantic Hierarchy, Direction, and Context Consistency in Standard Transformers
- GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them
- Separating Capability from Confidence: Grounded Dual-State Calibration for GRPO-Trained Medical Vision-Language Models
- CAVEAT: Recurrent Multimodal Diffusion Planning for Mapless Aerial Exploration
- CAROL: Context-Aware Online Learning for Fuzzer Scheduling
- Contextual Observer Grounding: Evaluating Situated Spatial Reasoning in Vision-Language Models
- Distributed Dexterous Manipulation with Spatially Conditioned Multi-Agent Transformers
- AstraMoE-SR: Trajectory-Guided Diffusion for Blind Satellite Jitter Deblurring and Super-Resolution
- CircuitLens: Reasoning Circuits as Data Selection Signals for Reinforcement Learning with Verifiable Rewards
- A Systematic Analysis of Automatic Differentiation versus Discretization-based Constraints for Physics-Informed PDE Solvers
- SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws
- The Role of Uncertainty in Assessing the Fairness of Machine Learning Models
- PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion
- DRIFT: Removing Diffusion Watermarks by Deflecting the Generative Trajectory
- ActionSplice: In-Flight Action Editing for Interactive World Models
- Non-Coherent Over-the-Air Federated Learning: Protocol, Convergence, and Device Scheduling
- Do Reviewers Still Reward Lexical Complexity? A Frozen-Rater Study of Preference Drift in 124K ICLR Reviews
- Temporal State Transport in Video Generation: Diagnosing and Correcting Spectral Imbalance
- Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing
- ZK-Trace: Certified Collusion Tracing with Zero-Knowledge Credentials for Federated GNSS Interference Monitoring
- A Closed-Form Estimator and Diagnostic Battery for Anchor-Judge Error Correlation, Under a Single-Common-Factor Model
- Projected Neural Differential Equations for Learning Constrained Dynamics
- Observability conditions for neural state-space models with eigenvalues and their roots of unity
- Modal Logic Neural Networks
- Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
- Diffusion-Inspired Reconfiguration of Transformers for Uncertainty Calibration
- Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning
- Marginal-Contribution Policy Gradients under Filtered Feedback for Multi-Agent LLMs
- $S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data
- RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
- Vocabulary-size-independent Convergence of Discrete Diffusion Models: adjoint equations induce the right space
- WMAttack: Automated Attack Search for Adversarial Evaluation of World-Model Agents
- Activation- and Influence-Aware Ranks (AIR): Function-Preserving SVD Compression for LLMs
- An AI4AI Framework for Visual Token Pruning
- CatchBench: When Can an Agent Failure Be Caught?
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- Clustering and Pruning in Causal Data Fusion
- Multi-Task Learning with Covariate-Overlap Regularization
- A robust and adaptive MPC formulation for Gaussian process models
- Accelerated Frank-Wolfe Algorithms: Complementarity Conditions and Sparsity
- Rotation-free Online Handwritten Character Recognition Using Linear Recurrent Units
- Token Encoding for Semantic Recovery
- State-Dependent Lyapunov Analysis of Rank-1 Matrix Factorization
- Masked Neural Detection for Run-Length-Limited Channel Coding in Molecular Communication
- Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and Tight Coordinatewise Rates
- Sharp Minimax Regret for Infinite-Memory Logistic Prediction
- Pooling and Drift in Delayed Bandits
- Look Before You Prompt, and After: Scaffolding Human-AI Collaboration in Software Tutorial Creation
- A Text Mining and Classification Approach for Analyzing Architecture Decision Records
- A Surrogate-based Approach for Fast Multi-objective Architectural Refactoring Optimization
- DJPlus: Generating minimal test suites for strong coverage criteria in graph models
- EventSpec: Defining and Detecting Event-Semantic Issues in Blockchain Ecosystems
- DREAMS: Modelling Support for Research into Engineering and Artistic Design
- An Empirical Study of the TianoCore Community
- A decade of rustls
- Extreme Server Side Rendering
- ID design and primary keys
- Changes to LLM pricing: GMICloud, Relace, StreamLake and Tencent
- No validator exists for llms.txt — here's the manual checklist
- Fable 5.1 Plays MMORPG Ultima Online For 2+ Hours
- Max20 to Max13 with Opus excluded from "all models"
- What actually happens after you hit enter in Claude Code — the whole path in one diagram
- Fable 5.1 vs GPT-6 Astra for 2D Sprites
- I was having trouble managing a lot of agents with just the terminal, so I built a "terminal development environment"
- Can an AI coding agent be locked out of modifying its own guardrail hooks? (OpenAI Codex CLI)
- Gave 6 AI models the same bug. Only 3 got it right.
- Catenary
- DeepSeek, Qwen, Zhipu AI Emerge Successively: PC Vendors Finally Get Long-Awaited AI Ammunition - 36 Kr
- Topology-induced Operators Reveal Complementary Graph Representations without Training
- Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes
- Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks
- Local and Global Stability in Performative Reinforcement Learning
- Behavioral Cloning Outperforms Entropy-Regularized RL: Critic-Driven Failure of Actor-Critic Methods on Adaptive Tumor Treatment
- LATS: Levy Adaptive Tree Sampling for Feedback-Driven Diverse Target Discovery
- Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
- AF-Mamba: Efficient Long-Term Signal Modeling for Early Prediction of Atrial Fibrillation Onset
- HyperTransfer: Understanding the Equivalence between Base Optimizer and Hyperball
- AI and TCAD for Inverse Design and Defect Discovery: From Simple Machine Learning to LLM
- Trust-But-Verify: Poisoning-Resilient Locally Private Graph Learning Protocols
- Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration
- Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration
- Translation of Black-Box Clinical Prediction Models into Standalone Transparent Nomograms: Temporal External Validation in Heart Transplantation
- ParetoTransport: Generative Optimization by Mass Transport Toward The Pareto Front
- Local gradient neural operator
- Decomposition-Guided Diffusion Language Models for Inertial Confinement Fusion Prediction
- Heat Field Signatures: From Point Clouds to Smooth Geometry
- Nystr\"om Attention Matches Full Attention for Cross-Sectional Stock Prediction
- SIM: Subspace Interaction-based Method for Token-Level Text Anomaly Detection
- Planet Labs' open satellite feed
- Paul Christiano joins OpenAI Foundation Board
- Shipt becomes the latest delivery app with an AI shopping assistant
- Instacart launches an AI grocery shopping assistant called Clementine
- IIns-VAE+: A Robust Transfer Learning Framework for Environmental Identification in Wireless Sensing
- Formation of structural attractors in neuromorphic systems
- Support Topology and Gradient Mixing in Sinkhorn Layers
- zScore-N: A Neural Network for On-Chain Wallet Reputation Scoring
- Robots Influencing Humans to Reveal their Goals during Collaboration and Competition
- TamilEOT: A Dataset and Model for Semantic End-of-Turn Detection in Tamil Telephone Speech
- Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education
- SIDE: Sensor Impersonation Detection at the Edge via Sequence Prediction
- Parameterized and Streaming Algorithms for Euclidean Fair $k$-Center Clustering
- MemCorr-DP: Counterfactual Correspondence Conditioning for a Diffusion Policy Guided by a Reference
- Companion-style QA Assistance in Ego-Vision
- AuthBench: A Large-Scale Multilingual Benchmark for Authorship Representation across Genres and Lengths
- From Synthetic Priors to Model Behavior: Structural Coverage in Tabular Foundation Models
- MSSP: Multi-Scale Spatially-Constrained Partition for Unsupervised Semantic Segmentation of 3D Point Clouds
- DPSF-Net: A Dual-Prior Spatial-Frequency Network for Real-World Remote Sensing Image Dehazing
- Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models
- Recompilation Is Not Enough: Test-Guided Decompiled-C Repair
- PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout
- Riemannian Optimization for Multi-Player Quantum Games on Product Unitary Manifolds
- TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables
- When Superpixels Fail on Documents: A Study of Segmentation for LIME Explanations
- Beyond the Matrix Sign: Quadratic Spectral Descent
- Let It Go or Learn to Self-Correct: Continuous Diffusion for Constrained Discrete Tasks
- Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
- Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives
- A concentration result for multilayer feedforward neural networks
- Adaptive Nonlinear Vector Autoregression: Robust Forecasting for Noisy Chaotic Time Series
- Knowledge-Guided Vision-Language Inference for Image-Based Urban Flood Depth Estimation
- Patent Representation Learning via Self-supervision
- Resolving sources of uncertainty in AI weather forecasting
- DL$^3$M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
- MedGround: Bridging the Evidence Gap in Medical Vision-Language Models with Verified Grounding Data
- FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
- Asking the Right Questions: Ontology-Grounded Interpretable Embeddings for Biomedical Text
- Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling
- FILT3R: Latent State Adaptive Kalman Filter for Streaming 3D Reconstruction
- UniQueR: Unified Query-based Feedforward 3D Reconstruction
- Aes3D: Aesthetic Assessment in 3D Gaussian Splatting
- An agentic framework for gravitational-wave counterpart association in the multi-messenger era
- EmoMind: Decoding Affective Captions from Human Brain fMRI
- More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
- A Fresh Look at Lamarckian Evolution and the Baldwin Effect
- Streaming LRAT Certificates into Lean Theorems
- Parameter-Free Dynamic Regret under Heavy-Tailed Noise
- Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models
- CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
- Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN
- StageWell: A Process-Aligned Chinese Corpus for Positive-Psychology Support Dialogue
- HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition
- GraphNOSE: A Graph Transformer in Olfaction
- A Multi-Source Ensemble Approach to Candidate Generation for Alternative Vacation Rental Property Recommendations
- Selective Posterior Margin Regularization for Forward-Corrected Classification
- QGB-W$k$NN: Quantum Granular-Ball Learning for Robust Classification
- Sparse Incident-Cluster Learning for 12-hour Port Flood Pre-warning in Digital-Twin Analytics
- Efficient Learning and Symmetry Discovery under Exact Invariances
- Kolmogorov--Arnold stability for discontinuous functions
- PCFlow: Physics-Conditioned Flow Matching for GPR B-Scan Image Synthesis
- Structured Extrema Errors in Classical Surrogates for Viscous Burgers: A Physics-Consistent Interpretation
- Automated Chest CT Protocol Selection via Large Language Model Derived Text Embeddings from Imaging Request Text
- HypLTSF: A Hyperbolic Geometric View of Multi-Scale Hierarchies for Long-Term Time Series Forecasting
- Stochastically Perturbed Weights: Ensembles from Deterministic Machine-Learning Weather Models
- Topological Fraud Detection in Latent Transaction Spaces
- AlphaRJM: Reward-Jump Memory for Stochastic Return-Guided Alpha Discovery
- HOPE: Heterophily-Aware Open-Set Node Classification with Pseudo-Extrapolation
- Length Generalization for Transformers via Compression
- Curriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics
- When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay
- ZetaDial: dialing net charge of protein binders at inference time for therapeutic developability
- Architectural and Regularization Components in Deep Learning Medical Image Registration: Systematic Ablation Study
- An Exploratory Study of Frequency-Aware Task Weighting for YOLOv8-Based Unified Driving Perception
- A Specialized Large Multimodal Model for Interpreting PET/CT in Head and Neck Cancer
- Characterizing Privacy Risks of Quantum Machine Learning with Emergent Quantum-Native Access
- RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
- SLA-Safe Energy Control for AI-Native NG-RAN Using Stability-Aware Constrained PPO
- FALCON-S: Fixed-wing ground-effect Aerodynamics Simulator and Flight Control Learning Suite
- SolarBench: A global solar energy nowcasting benchmark
- One Shared LoRA Weight for MRI Reconstruction across Acceleration Factors
- Large-Scale Pretraining for Improving Deep Learning-Based Geometric Distortion Correction of Diffusion-Weighted Imaging
- Introductory Notes on Learning$^2$
- BinauralVAE: Spatial Audio Reconstruction For World Models
- Beyond Task Success: Stage-Wise Reliability of World Model Planning under Sensing Degradation
- Graph neural networks and the energetic cavity method for combinatorial optimization
- SeisBench DAS: A machine learning framework for Distributed Acoustic Sensing
- Cross-modal learning for SAR target recognition using optical vision foundation models
- JEDI: JEPA-to-Edge Distillation for Efficient Cropland Segmentation from Satellite Imagery
- MI-PEFT: Mixture-of-Experts Integrated Parameter-Efficient Fine-Tuning Protein Language Models Improves Acidophilic Proteins Classification
- Speed Limit for Information Acquisition in Stochastic Learning Dynamics
- Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTM
- Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules
- FedGenSC: Federated Generative Semantic Communication with Channel-Aware Adaptation
- Silver Rate Is (Almost) Optimal for Gradient Descent Acceleration
- Understanding Uncertainty Sampling via Equivalent Loss
- TSMini: A Simple Yet Highly Effective Trajectory Similarity Learning Model
- Two-dimensional Taxonomy for N-ary Knowledge Representation Learning Methods
- Learning Latent Graph Geometry via Fixed-Point Schr\"odinger-Type Activation: A Theoretical Study
- CDFlow: Building Invertible Layers with Circulant and Diagonal Matrices
- Temporal Kolmogorov-Arnold Networks (T-KAN) for High-Frequency Limit Order Book Forecasting: Efficiency, Interpretability, and Alpha Decay
- Continual Policy Consolidation for Lifelong Robot Learning
- Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
- FuseDiff: Symmetry-Preserving Joint Diffusion for Dual-Target Structure-Based Drug Design
- EcoFair: Energy-Efficient Inference Routing for Edge AI under Data Degradation
- PMCTS: Principled Parallelized Inference Time Scaling with Particle Monte Carlo Tree Search
- Characterizing Privacy-Audit Alignment in Behavioral Audit of Machine Unlearning
- When Does Activation Steering Change What a Model Computes From?
- Can an AI Assistant Really Forget? Auditable Deletion from Addressable Memory
- TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series
- Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra
- A Storage-Retrieval Gap in Parametric Knowledge Graph Memory
- D-TAIA: Domain-Aware LLM Adaptation for Multi-Task Predictive Process Monitoring
- InKAN: B-Spline KANs via Truncated Power Form
- Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications
- MILAAP: Mobile Link Allocation via Attention-based Prediction
- The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
- Solving the Offline and Online Min-Max Problem of Non-smooth Submodular-Concave Functions: A Zeroth-Order Approach
- Stability of the Monge Map in Semi-Dual Optimal Transport
- Expressivity of Contradiction Graphs
- PLC-Bin2Src: Retrieving Corresponding Structured Text Source Files for PLC Binaries
- The State of Allocators in 2026 - 6 Months Later
- The state of European cloud providers in 2026
- Does anyone here actually let Claude (or any agent) touch their email?
- Built a tiny touch display for managing Claude Code, Chibi Deck.
- I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.)
- DeepSeek Initiates Pursuit of Ultimate "Intelligence Density" – Redefining AI Innovation & Performance - 36 Kr
- 巨大オンラインゲーム「EVE Online」。ゲームを支える240万行のPython 2.7のコードをPython 3へ移行すると発表
- Sparse Oblique Rule Boosting for Simpler Additive Rule Ensembles
- Bi-HYCO: Bi-Objective Cooperative Learning for PDE Parameter Identification under Fragmented Observations
- Constrained Bayesian Optimization for Hierarchical Federated Learning in IoT Networks for Plant Disease Classification
- PPIM: Pennes Physics-Informed Mamba for Heat-Source-Conditioned 3D Bioheat Simulation
- HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care
- Improving Multivariate Time Series Classification with Class-Wise Training and Model Aggregation
- Attributing Cohen's d: Training Data Attribution for Disease-Related Effects in Normative Age Biomarkers
- Guiding Worker Self-Selection in Crowdsourcing Contests: An LLM-Augmented Algorithmic Approach
- Semi-Supervised Learning under Spatially Biased Sampling
- Not All Variables Agree: Reliability-Aware Variable-Wise Gradient Surgery for Multivariate Time-Series Forecasting
- Why shared attention vectors fail: a case for outcome-indexed tuning
- BAFF: Bid-Aware Filter Family for Mitigating Training Data Interference in RTB A/B Tests
- Condition aware learning enables robust prediction of oligonucleotide melting behavior across diverse chemistries and assay conditions
- Infrastructure-based Monocular 3D Vehicle Localization Framework with Experimental Validation
- Rollcast: Proper-Score Gated Rolling Anchors for Adaptive Probabilistic Time-Series Forecasting
- Diagonal Attenuation: A Finite-Sample Correction for PCA
- Polarity-Asymmetric Structural Calibration for Link Sign Prediction
- Spatial Attention Supervision for Defect Localization: Exploiting Ground-Truth Masks as Training Signal in Diffusion-Augmented Defect Detection
- Beyond Worst-Case Coreset Bounds for $k$-Clustering via Determinantal Sampling
- Stochastic Nonconvex Bilevel Optimization: Improved Rates Without Rare-Visit Assumption
- Accelerated High-Accuracy Sampling from a Warm Start via the Proximal Bouncy Particle Sampler
- Smoothed Picard Hamiltonian Monte Carlo
- Structural Entropy-Driven Graph Diffusion Generation for One-Shot Federated Graph Learning
- Role-Specific Predictive Geometries for Nonstationary Multivariate Graph-Signal Forecasting
- Constitutive State-Space Modeling of Path-Dependent Plasticity: A Resolution-Consistent and Parallelizable Computational Framework
- Impact of canny edge detection preprocessing on performance of machine learning models for Parkinson's disease classification
- Statistical versus machine learning-based spatial interpolation of post-processed ensemble weather forecasts
- DeepSeek seeks 150 engineers in hiring spree to overhaul back end systems - South China Morning Post
- Claude, change the "Add to Cart" button to blue
- Deep belief networks are exact
- Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement
- Explainable Temporal Attention-based Defect Detection For Fillet Joints in Real-Time Gas Metal Arc Welding Based on Multi-modal Data
- RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
- When Can One Obtain Certificates of Optimality Using Positivstellensaetze?
- Time-Varying Data as Sheaves: an Invitation to Narratives
- A Generalization of Amari's Bayesian Duality
- Neural Symbollic Regression Using Deep Learning and Sparse Modelling
- TC-Next: Zero-Shot Multimodal Cyclone Forecasting
- Diffusion models for eye-gaze trajectory generation using position and velocity representations
- XAI-SDN: An Explainable Entropy-Guided Machine Learning Framework for Real-Time DDoS Detection in Software Defined Networks
- SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation
- PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
- Recovering Weak Signals with Normalizing Flows
- OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
- Discovering Translation-Worthy Languages with E-Values
- Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model
- AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories
- Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
- Re-calibrated Contrastive Loss for Transformation-Aware Prompt Conditioning in Vision-Language Models
- Solution for UCF UrbanTwin LUMPI Track: Sim-to-Real Urban LiDAR 3D Object Detection
- Thermodynamic Cyclic Processes with Markov Samplers in Bayesian Inference
- NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis
- Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
- From Proxies to Fields: Spatiotemporal Reconstruction of Global Radiation from Sparse Sensor Sequences
- MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment Using UAV Imagery
- S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
- AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification
- SAGE: Shape-Adapting Gated Experts for Adaptive Histopathology Image Segmentation
- Tracing Mathematical Proficiency Through Problem-Solving Processes
- Diffusion Model in Latent Space for Medical Image Segmentation Task
- Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs
- US-JEPA: A Joint Embedding Predictive Architecture for Ultrasound
- Spatiotemporal Heterogeneity of AI-Driven Traffic Flow Patterns and Land Use Interaction: A GeoAI-Based Analysis of Multimodal Urban Mobility
- Field-level weak lensing cosmology with $60$ simulations using multifidelity simulation-based inference
- TRNet: Learning with Topographic Priors for VHR Paddy Rice Mapping
- Decoupled Temporal Encoding for Generative Recommendation
- Deep Learning-Based Multi-User Communication Design for Dense IoT Networks: Interference-Aware Finite-Blocklength Communication and Preliminary MIMO Extensions
- Multi-granularity Adaptive Hypergraph Representation Learning via Granular-ball
- Scaling Optimal Classification Trees via Adaptive Feature and Sample Reduction
- A budget-dependent crossover between coverage- and response-based training-set selection for machine-learned interatomic potentials
- A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret
- A dictionary learning framework for graphs via filters and optimal transport
- Granular-Ball Quantum Clustering for Resource-Efficient and Robust Learning
- IXPLORE: Bounded Ideal Point Estimation with Grid-Based Uncertainty Quantification
- Minimizing the Effect of Sleep Deprivation in the Forward-Forward Algorithm
- Adaptively Incorporating Directional Hints into Zeroth-Order Optimization
- EMBLEM: Enhancing Multi-script Table Detection through Masking
- Geographically Regularized AUC-Maximizing Personalized Federated Learning
- Certified Topological Interaction in Neural Representations: Class Disentanglement Is Mostly Pairwise
- Multi-Level-Set-Based Physics-Driven Neural Network to Solve 3-D Inverse Scattering Problems
- PAC-Bayesian Bounds for Learning Partially Observed Stochastic Linear Time-Invariant State-Space Systems with Inputs and Sub-Gaussian Noise
- Physics-Informed Deep Learning for False Ventricular Tachycardia Alarm Reduction in the ICU
- Multi-Task Learning for Sparsely-Labeled Time Series: A Case Study on Cold-Hardiness Modeling
- Nearly Tight Rademacher Bounds for Sparsely Activated Neural Networks
- Novel hybrid protein scaffold gap filling using weighted machine learning ensemble, beam search, and mass-constrained reranking
- Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures
- Asymptotically-informed neural networks for Black-Scholes implied volatility computation
- Contrastive Knowledge Distillation for Anomaly Detection in Multi-Illumination/Focus Display Images
- Physics-Informed Neural Networks for Depth-Averaged Granular Avalanche Dynamics on Curved Topography
- A Nuclear-Norm Lower Bound for Dithered Scalar Quantization of Matrix Products
- Tight Lower Bounds for State Tomography with Limited Entanglement
- Gaussian Linear Functional Manifold Method for Massive Point Cloud Data
- AAMBERS-UAV: Acquisition-Aware Multimodal Backbone Evaluation and Ranking for UAV Weedy Rice Segmentation
- Physical policy gradient theorem for in situ stochastic-adjoint training
- Functional Attentive Interpretable Regression
- Causal DAG Identification for Count Data via Poisson Thinning Structural Equation Models
- Recovering linear images of sparse signals from indirect observations
- Fast PAC Global Optimization via Restarted Langevin: Exploration, Exploitation, and Degenerate Cooling
- Robust conditional dimension reduction for dissimilarity data
- Hierarchical Fourier Approximation for Variational Quantum Distribution Learning
- Deep learning from the crowd Fundamentals of morphological galaxy classification
- Likelihood-Based Unsupervised Anomaly Detection in CMS Dijet Events
- Large Classification-Risk-Optional Label Acquisition
- Detect Anything in Graphic Design: Element-Level Rewards for Autoregressive Detection
- Heat Kernel Textures: the Geodesic Gaussians That Do Not Splat
- A Sub-4 Approximation for Fair $k$-Means
- Flexible Motion Generation from Language and Style References
- MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States
- A Gradient-based yet Spike-Timing-Dependent Solution to the Feedback Learning Problem in Neural Microcircuits
- How to Make the Gradient Mapping Small for Constrained Stochastic Min-Max Problems and Beyond
- The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives
- Learning to build covering structures with continuous adjustments
- ONE CYLinder: A Benchmark for Graph-Based Surrogate Modeling of Unsteady Bluff-Body Flows
- High-dimensional Linear Bandits with Knapsacks
- Deep Learning Approach to Bearing and Induction Motor Fault Diagnosis via Data Fusion
- Deep Learning to Automate Parameter Extraction and Model Fitting of Two-Dimensional Transistors
- Comparison of D-Wave Quantum Annealing and Gibbs Monte Carlo for Sampling from a Probability Distribution of a Restricted Boltzmann Machine
- Electricity Price Forecasting: Bridging Linear Models, Neural Networks and Online Learning
- Multi-Fidelity Physics-Informed Neural Networks with Bayesian Uncertainty Quantification and Adaptive Residual Learning for Efficient Solution of Parametric Partial Differential Equations
- Text Has Curvature
- Towards Near-Real-Time Telemetry-Aware Routing with Neural Routing Algorithms
- Learning Without Adversarial Training: A Physics-Informed Neural Network for Secure Power System State Estimation under False Data Injection Attacks
- Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge
- Federated Foundation Models over Vehicular Networks
- Neural Field Tokenizations with Hierarchy and Spatial Locality Priors
- When to Align, When to Predict: A Phase Diagram for Multimodal Learning
- Monotonic Kolmogorov-Arnold Networks: A Theoretical and Empirical Study of Monotonicity as an Inductive Bias
- Multistage Defer Trees for Hybrid Interpretability: If at First You Can't Succeed, Tree Again
- What a Deletion Certificate Covers, and Where It Expires: Auditable Removal from a Support-Vector Memory
- AI-Augmented Adaptive Digital Twin Modeling for Brain Tumor Evolution Prediction and Treatment Scheduling
- Information-Theoretically Secure Aggregation for Lightweight Federated Learning: Resilient to Dropouts and Adversaries
- ArborEnum: Decision Tree Rashomon Sets over Continuous Features
- Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training
- Euclidean Fourier Neural Operators
- PathGuide: Dynamic Classifier-Free Guidance via On-Policy Transport Alignment
- Generative artificial intelligence for reliable mechanistic reasoning for corrosion
- Spectral-Target Physical Latent Structuring for JEPA-Style World Models
- Generalized infinite dimensional Alpha-Procrustes based geometries
- An Empirical Study on the Impact of Change Granularity in Refactoring Detection
- Solaris Turnstiles
- Rust: When Empty Isn't Bottom
- Why Function Arguments Are Not Function Colors
- Every Millisecond Counts
- Robo-Advisory Platform: Essential AI Wealth Access
- Verification & Security
- Changes to LLM pricing: Baidu
- I started using Codex to learn DevOps… now I’m wondering what exactly I’m learning 😂
- GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
- Solved: LLM inference on Windows was 2–3x slower when the server window wasn't focused
- Is anyone working on conversation compaction?
- Indianapolis man accused of creating sexual videos of children with AI - Fox 59
- Model-Adaptive and Risk-Constrained Frequency Hopping Against Predictive Jammers
- CLUES-WEASEL: No additional clues required to choose your time series clustering algorithm
- A Machine Learning Framework for Predicting Restaurant Food Waste to Support Sustainable Food Management
- Learning Kernels by Alignment for Multiclass Bayes Classification
- Learning Adaptive SED for heterogeneous load balancing
- Inferring Urban Mobility Interactions from Aggregated Dynamics
- Emergent Charging Coordination in Electric Delivery Fleets
- The Work Now Within Reach
- Quantile-Led Feature Extraction for Multi-Horizon Predictive Maintenance in Industrial Manufacturing Systems
- Three Types of Negation of Triple and its Elements and an Extension of Triple
- Data-driven rational function neural networks: a new method for generating analytical models of rock physics
- Convergence issues in Relational Concept Analysis based on AOC-posets
- ViT3Flow: A Test-Time Training Transformer MeanFlow for Postoperative Radiograph Synthesis in Scoliosis
- Do Quantum AIs Dream in Paths? Path-Integral Slow Thinking through Grover Interference
- What Does Animal Re-Identification Learn? Linear Biological Concepts and Their Origins in Visual Representations
- Calendar-SPCA: Interpretable Representation Learning for Multi-Periodic Electricity Consumption Profiles
- Programmable Cellular Automata
- AGSA-Net: Abundance-Guided Self-Attention Network for Spectral Unmixing-Aware Hyperspectral Remote Sensing Image Classification
- Layer-Wise Gate-Controlled Prompt Truncation in a Multimodal Chest X-Ray Classifier
- SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives
- TD-STGT: A Spatio-Temporal Graph Transformer for Mobile Traffic Demand Forecasting
- FSAN: Flow State Attention Network for Aerodynamic Prediction
- Attention-Enhanced Deep Features with Heterogeneous Ensemble Learning for Glaucoma Detection
- AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering
- Emo-DVS: A Multimodal Benchmark for Privacy-Aware Emotion Recognition with Event Cameras
- AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation
- CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records
- Tensor network representations of discrete maximum entropy distributions via mean polytopes
- D3ARC: Time-Critical Distributed Disaster Detection for Asynchronous Cooperative Multi-Robot Systems
- Generation of Vectorized Maps Beyond Vehicle View
- Microcanonical Hamiltonian Monte Carlo and the Helmholtz Theorem
- Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials
- Fusing Sequence Motifs and Pan-Genomic Features: Antimicrobial Resistance Prediction using an Explainable Lightweight 1D CNN-XGBoost Ensemble
- Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular Maximization
- Nonlinear elliptic homogenization with the parametric Deep Ritz method
- Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study
- Sector-Mean: Deterministic Initialization of K-Means Centroids via Angular Sector Partitioning
- No-Regret Mixing of LRU and LFU with Optimal Switching Cost
- Online Signature Verification Using Augmented Path Signature and T-Mamba
- Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts
- Iterative Audio Separation with Mixture Consistency via MIMO Model Extension
- Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution
- A Quantitative Evaluation Framework for Temporal Explainability in Echocardiographic Video Segmentation
- Bayesian Matrix-Valued Graphs for Context-Dependent Multivariate Relationships
- A Transformer-Based Delta Expression Encoder for Psilocybin Transcriptional Response: Architecture, Representations, and Biological Validation
- Selective boundary condition reduction via learned error gating
- Flexible Spectral-Normalized Neural Gaussian Process for Dynamic Aperture Prediction
- A Note on Scaling in Randomly Rotated Quantization and Its Connection to the CDEF +1 Pythagorean Relation
- High-Magnetization Sampling at Low Temperatures: Ising Models and Bayesian Sparse Linear Regression
- Fitting and Learning Basis-Restricted Propositional Formulas
- Latent class analysis by regularized spectral clustering
- On the Effectiveness of the z-Transform Method in Quadratic Optimization
- Detecting and explaining clinical-omics inconsistencies to improve patient cohort stratification: an application to Parkinson's disease
- Best-of-Both Worlds for linear contextual bandits with paid observations
- Minimum distance classification for nonlinear dynamical systems
- Echo State Networks for Time Series Forecasting: Hyperparameter Sweep and Benchmarking
- Feedback Control for Multi-Objective Graph Self-Supervision
- Learning functional components of PDEs from data using neural networks
- Temporal Consistency Improves Generalization in Contextual Offline Meta Reinforcement Learning
- Generating from Discrete Distributions Using Diffusions: Insights from Random Constraint Satisfaction Problems
- Subspace Optimization for Backpropagation-Free Continual Test-Time Adaptation
- Conformalized Quantum DeepONet Ensembles: Towards Scalable Operator Learning with Distribution-Free Guarantees
- Learning Polyhedral Conformal Sets for Robust Optimization
- PUID: A Personalized Deconfounding Framework for Recommender Systems under Hidden Confounding
- PCA-Enhanced Adaptive NVAR Framework for High-Resolution Sea Surface Temperature Forecasting in the East Sea
- VegSim: A Geospatial World Model for Scenario-Conditioned Vegetation Simulation
- Probing Chemical Language Models: Effects of Pre-training and Fine-tuning
- Masked Generative-Contrastive Representation Learning for Cross-Dataset EEG-Based Emotion Recognition
- When Does Reward Teach State? A Hidden-Automaton Instrument and a Group-Language Warning Signal
- On the Potential of Graph Neural Networks as Metamodels for Supply Chain Optimization: Dataset, Architectures, and Directions
- Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering
- An Identifiability Theory of Masked Prediction: Mode Blindness and Mask Schedules
- Noise in Diffusion Models Is a Learnable Input
- Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group
- Advancing Open and Reproducible Relational Learning: RelArena-$\alpha$, TabPFN-Rel and RPI
- Coordination on a Budget: Federated Active Learning with Few Labels
- Temporal Memory-Aware Online Test-Time Adaptation on Dynamic Graphs
- Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling
- Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions
- Human-Robot Interaction and Perceived Irrationality: A Study of Trust Dynamics and Error Acknowledgment
- The purpose of DNS is to spread scams
- Reverse engineering my e-scooter and rewriting the firmware in rust
- I changed my license to EUPL
- 10 AI Website Builders I Tested So You Can Skip the Trial and Error
- Por qué el 80% de los profes que prueban ChatGPT lo abandonan en 2 semanas
- Changes to LLM pricing: Alibaba, Baidu and StreamLake
- The leftover
- Astra is a fantastic for coding
- Now this is a serious local machine
- Qwen/Qwen-Drive-1.0-4B · Hugging Face
- Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits
- Type.com
- 'All the AI is down': ChatGPT, Claude, and Grok crash at once, sparking relief online - The Cool Down
- 5 Ways I Access Coding Models for Free - KDnuggets
- Forecasting the Winner of a Live Tennis Match
- Roame (YC S23) Is Hiring Viral Content Editor
- Modus Tollens and Counterfactuals and Counterfactual Reasoning Based on Three Types of Negation
- Subject-Relative Micro-Motion and Sleep Dynamics for Near-Infrared Video Sleep Staging
- Full-Page Optical Music Recognition of Handwritten Monophonic Scores
- Multiple Myeloma Lesion Segmentation on Whole-Body Diffusion-Weighted Imaging via Efficient Anatomical Anticipation and Multimodal Confirmation
- Recovering topological information of light by topological learning
- Deep Barycentric Regression for Optimal Transport Map Estimation and its Statistical Optimality
- When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization
- Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning
- FedRAW: Preserving Rare-Label Influence in Asynchronous Federated Learning
- Federated Binary Gating with Server-Side Vision-Language Inference for Surveillance Anomaly Classification
- Feature Reconfiguration With Visual Prior for Medical Lesion Segmentation
- Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids, and Full-Bandit Learning
- Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach
- Anchored Regularized Direct Least Squares (ARDLS): Integrating Established Prioritization Operators for Priority Elicitation in the Analytic Hierarchy Process
- Improving Randomized Metric Distortion to 2.1441
- Spatial Feature-wise Linear Modulation (SpFiLM) for Contrast Agent-Aware Brain Parcellation
- Sub-6 GHz Over-the-Air AMC via Curriculum Fine-Tuned CNN-Transformers
- Optimal Slice-Adaptive Tuning of Hybrid Slice Sampling
- Distribution-free inference on the number of changepoints
- Inclusive electron-nucleus cross section models from domain adaptation
- Non-Adaptive 1-Bit Mean Estimation: Minimax Rates and the Sample-Interval Tradeoff
- Optimal estimation for Functional Linear Regression with Noisy Discretized Data
- From Human Labels to Literature: Semi-Supervised Learning of NMR Chemical Shifts at Scale
- Low-Rank Plus Sparse Matrix Transfer Learning under Growing Representations and Ambient Dimensions
- Improved Dimension Dependence for Bandit Convex Optimization with Gradient Variations
- A Thermodynamic Theory of Learning Part II: History-Dependent Reachability and Continual Learning
- Constant-Stepsize Stochastic Approximation: Finite-Time Convergence, Gaussian Approximation, and Tail Bounds
- Mixing Makes Markovian Contexts Cheap for Linear Bandits
- The Geometry of Polynomial Group Convolutional Neural Networks
- Domain-Aware Hybrid Quantum Learning via Correlation-Guided Circuit Design for Crime Pattern Analytics
- Near-Floor Geometry Is Generic: Leverage Dispersion in Trained Overcomplete Codes
- Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
- Quantum Hierarchical Reinforcement Learning via Variational Quantum Circuits
- Reactive Flux Matching: Mechanism Discovery and Adaptive Sampling of Rare Events
- Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
- Learning Subgroup Relations Using Siamese Graph Neural Networks
- Online Convex Optimization with Dueling Feedback
- K\"ahler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold
- Across-Design Uncertainty in Short Pricing Panels: Inference and Identification
- From Relaxed Indexability to Exact Indexability: A $t$-Step Approach for Partially Observable Restless Bandits
- A Geometric Phase Boundary for Volume-Sampled Linear Readouts
- Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training
- Directed mixed membership stochastic blockmodel
- Tempora-Fusion: Time-Lock Puzzle with Efficient Verifiable Homomorphic Linear Combination
- I Don’t Want to Interact With Stochastic Parrots
- software for humans
- Why is this robot slowly dropping the object?
- Unlock Turkey’s Digital Transformation: AI-Powered Solutions for Businesses & Education
- Unlock Vietnam’s Potential: How AI & Smart Automation Can Drive Your Business & Skills
- Zenoti Consulting Services for Connected Business Workflows
- I made a virtual lounge for vibecoders to hang out while claude code is running.
- what did you stop using claude for after trying it?
- Playlist Atlas built with Fable 5.1
- What is the best and most balanced effort to code with Fable 5 on the $100 plan?
- Too used to Claude Code to switch?
- Mention if a "new model" is a finetune
- Qwen3.8-27B-Uncensored-Genesis-V1-GGUF
- Best Open source TTS right now for narration?
- What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?
- Frigade Assist API
- DeepSeek Overhauls Interview Process: New Questions Stun Even Fresh Graduate ACM Gold Medalists - 36 Kr
- Grok Bot Was Built in Just Four Weeks - analyticsindiamag.com
- A Statistical and Machine Learning Framework for Quantifying Offensive Impact in Professional Box Lacrosse
- Claude or GPT for academic workflows?
- Codex stuck on commands – anyone else?
- What AI subscription should I switch to?
- Codex $100 or grok $100 for langgraph/langchain development?
- Server rebuild to custom loop. 2x RTX Titans 24gb, 1x 22gb 2080ti | T: 70GB VRAM.
- new Nex model
- Qwen is working better for me than ChatGPT – what’s the best way to use it with ZIP/code projects?
- bonds
- xAI denied injunction against Minnesota AI-nudification law - Minnesota Lawyer
- Grok and FSD Merging One Day? What Owners Should Know - BASENOR - Tesla Accessories
- Recreating a 70-year love story frame by frame
- Analysis of Respiratory Sinus Arrhythmia with Neural Networks
- Masking Radar Cognition under Adversarial Surveillance: A Distributional Privacy Framework
- Pre-Whitening and BCJR Posterior Distillation for Bi-LSTM Detection in Faster-than-Nyquist Signaling
- A Universal Reproducing Kernel Hilbert Space from Polynomial Alignment and IMQ Distance
- Linear and Quadratic Discriminant Analysis: Tutorial
- Switching Password Managers in 2026
- Analysing 2048 on a 3×3 board
- You’re right, and it’s worse than I thought
- Load-bearing humor
- Is my current workflow sufficient?
- Has anyone ever seen this before?
- Astra should be removed from Pro.
- How do I make websites generated by Codex look better
- Quoting Terence Tao
- Turn one AI prompt into multiple answers with this $55.30 lifetime subscription - PCWorld
- How a 98-Year-Old Uses Tesla FSD and Grok Every Day - BASENOR - Tesla Accessories
- 「ひかりでんわ」を作って学ぶ、子ども向け通信実験教室をやった話
- Great idea, Claude!
- Does anyone have a better way to connect Claude to multiple Gmail profiles?
- Why the hell is LM Studio making LM Studio so difficult to download?
- Don't let FOMO win if you're interested in local llm from a hobby/learning aspect
- Basedash in Español Français & Português
- Trancy Air
- XRP Price Prediction: We Asked Grok Where XRP Ends September - 24/7 Wall St.
- Elon Musk Grok AI Predicts $250K Bitcoin Price by 2027 - Cryptonews
- Is SpaceX a Top Artificial Intelligence (AI) Stock Pick in September? - The Globe and Mail
- China Says US Accusations of AI Theft Are ‘Groundless’ - ASHARQ AL-AWSAT English
- A solution to the Erd\H{o}s Problem #1040
- How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits
- Multi-label versus multi-class classification of blood cells and their aggregates in microfluidic channels
- The Art of Hierarchical Competing Patterns: Gaussian Process Optimization of Hyphenation
- Closed-Form of the Local Galactic Potential and Stellar Distribution Function from Gaia DR3
- How to build a f**king printer
- An AI's Completely Ordinary Day (A True Story)
- Opus Simulator
- We need to talk about this icon in "Design"
- I dont understand Skills and need help
- Grok simulates the final outcome of the New England Patriots’ 2026 season ahead of Week 1 Super Bowl rematch against Seahawks - A to Z Sports
- ASEAN Must Be Cautious About Growing Chinese Trade And Investments – OpEd - Eurasia Review
- Grok simulates the final outcome of the Philadelphia Eagles' 2026 season ahead of Week 1 against the Washington Commanders - A to Z Sports
- iPhone Duo
- Google DeepMind Maps 9 Billion Possible DNA Variants
- Tailwind Labs is joining Shopify
- What do Visa and Mastercard do? An intro to card networks
- GNU Radio in the browser
- Apple Watch Ultra 4
- Understanding the recent DDoS attack against Read the Docs
- Apple Introduces AirPods 5
- Why Emacs Consult async searches feel slow and how to speed them up
- Generating the P3 Tiling
- We accidentally built a synthetic cell factory
- Apple Watch Series 12
- No Man's Sky Cosmos
- Get ready for the game with new football features in Search
- Why Rider and ReSharper Were Slow to Start, and How Microsoft Helped Fix the Problem
- Join our live webinars: Migrating from Atlassian to YouTrack
- The Evolution of WSL Support in JetBrains IDEs
- dotInsights | September 2026
- Debug Past an HTTP 403 Without Breaking Spring Security
- 6 Benefits of Sandbox Environments (and How Docker Sandboxes Delivers Them)
- Besxar is building an orbital semiconductor factory, one SpaceX rocket at a time
- GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- Suno launches v6 music models built with Warner, BMG, and Believe
- Information entropy as an anthropomorphic concept
- Apple Event for September 9th, 2026
- Ass Auction
- DuckFightClub
- Illustration shows Deepseek logo and "Artificial Intelligence AI" words - The Daily News | Texas' Oldest Newspaper
- CLAUDE.md を憲法にして、AI の完了報告を受け口で機械検査した記録 (いいね相当スコア: 0)
- 外部依存ゼロの日本語意味理解エンジン KotobaCore を 1.0 にした — 「誰が何にどう感じたか」まで構造化する (いいね相当スコア: 0)
- Mac Studio M3 Ultra 96GBでローカルLLMを検証してみた (いいね相当スコア: 0)
- 30B級ローカルLLM、現場で使うならどれ?Qwen3.8・Muse Glimmer・Gemma4を比較【コーディング編】 (いいね相当スコア: 0)
- Three.js × Claude Visionで「商品が売り場でどれだけ目立つか」をシミュレーションしてみた (いいね相当スコア: 取得失敗)
- Kimi K3の1M文脈を支えるKimi Delta Attentionという線形注意 (いいね相当スコア: 1)
- RAGの精度は「データ」で決まる — Contextual Retrievalを実測したら1位正解率が36%→100%になった (いいね相当スコア: 3)
- 「求人票を1件ずつチェックする」案件探しをClaudeで自動化するツールを作った (いいね相当スコア: 1)
- 日本語RAGのテストケース設計 — 「文書にない」を正解にする (いいね相当スコア: 0)
- 多様な視点が暗記に勝つ──補助ビューがLLMの事前学習を加速する仕組み (いいね相当スコア: 0)
- LLMの構造的限界から考える、AIエージェントの設計原則 (いいね相当スコア: 0)
- 2026年4月以降、Einoはどこまでエージェント基盤になったのか (いいね相当スコア: 3)
- 検索ヒット率をLLMで底上げしつつコストを抑える『安価モデル事前スクリーニング』設計 (いいね相当スコア: 0)
- 1971 年のロボットは、どうやって計画を立てていたか — STRIPS と、いまのエージェント (いいね相当スコア: 0)
- LLMのトークン効率化で気をつけたいことまとめ (いいね相当スコア: 20)
- 【図解】TPかPPか問題:社内ローカルLLMを複数GPUに載せる判断軸 (いいね相当スコア: 0)
- 勝手に最新へ移るLLMエイリアスを固定IDで縛る判断基準 (いいね相当スコア: 0)
- AIエージェントにセキュリティ要件をどこまで任せられるか — ASVSで仕分けたら3種類しかなかった (いいね相当スコア: 6)
- 顧客に会う前夜、オレは自分が結んだ契約の中身を知らなかった——自分で考えていない案件は、驚くほど理解できない (いいね相当スコア: 6)
- 2026-08-24 今日の技術トレンド (いいね相当スコア: 0)
- 日本語RAGは「入れれば賢くなる」にならなかった ― KotobaCore開発秘話(1/8) (いいね相当スコア: 0)
- 知能の本質は"確率予測"なのか——ハルシネーションから道具的収束、感情の正体まで (いいね相当スコア: 1)
- GPT-2-likeからQwen2-likeへの実験:第4回 MHAをGQAに変える (いいね相当スコア: 2)
- VLM量子化はLLMだけ見ない:ViT・Connector・LLMのbit配分を考える (いいね相当スコア: 0)
- JEPXの価格を時系列基盤モデルで予測する:越えるべき壁は「昨日のコピー」だった(前編) (いいね相当スコア: 0)
- 推定量の有効性と一様最小分散性の違い (いいね相当スコア: 0)
- Web Sustainability Guidelinesを読む会 #5 (いいね相当スコア: 2)
- 「相関の符号が反転した」のは市場ではなく装置だった — 時系列観測で最も見落としやすいメタデータの話 (いいね相当スコア: 0)
- SharQの式を追う:FP4と4:8 activation sparsityの2経路設計 (いいね相当スコア: 0)
- 予測精度を1%上げると、いくらの価値があるのか (いいね相当スコア: 0)
- LLMは時系列予測ができるのか (いいね相当スコア: 0)
- AIエージェントの評価(Eval)について考えてみた──ベンチマークが当てにならない理由 (いいね相当スコア: 0)
- 「文章を続けるモデル」から紐解くTransformerの仕組み (いいね相当スコア: 0)
- AIエージェントにファイルを消される・課金が止まらない・秘密鍵が漏れる|暴走の原理と4層の対策を調べてみた (いいね相当スコア: 6)
- 【2026年8月】Claude Fable 5.1がリリース!SWE-bench Pro首位奪還・Swarm Mode・API変更点まとめ (いいね相当スコア: 0)
- 【2026年9月】GPT-6 Astraがリリース!4Mコンテキスト・Verbosity制御・API変更点まとめ (いいね相当スコア: 1)
- 「Reply with exactly: t1」が$13.61だった — GPT-6 Astraの週枠が溶ける場所の実測 (いいね相当スコア: 0)
- 「ノーコードだから早い」は思い込みだった、低コードMLの導入に平均4.5ヶ月かかる件 (いいね相当スコア: 0)
- 【話者分離】業界標準のPyannoteと同水準の自作アルゴリズムを実装した話 (いいね相当スコア: 0)
- 【技術解説】【完全ガイド】Pythonによる自動売買システムのバックテストと実践コード (いいね相当スコア: 0)
- 学習後に話速は変えられない — パラメータは受け取るのに、無視されていた (いいね相当スコア: 0)
- 月10万円で組織が使えるLLM基盤をどこまで作れるか|その3:使える性能を考える (いいね相当スコア: 取得失敗)
- そのAIレビュー、信じて大丈夫?──「いいですね」に潜む迎合とプロンプト設計 (いいね相当スコア: 取得失敗)
- ローカルLLMを9Bから27Bにしたら正答が13/20→19/20。ただし新しい世代の27Bは17/20に落ちた (いいね相当スコア: 取得失敗)
- 第七話「炉苦夢経という道もある」 (いいね相当スコア: 取得失敗)
- 広島AIプロセス報告枠組み2.0 / 自主開示の実質を誰が検証するか 雑感 (いいね相当スコア: 取得失敗)
- AIエージェントの忠実義務 / 開示規制の限界と比較推奨販売の経験 雑感 (いいね相当スコア: 取得失敗)
- 損害なきインシデント / エージェント逸脱事案と報告閾値の設計 雑感 (いいね相当スコア: 取得失敗)
- AIエージェントの統制が届く範囲と責任が及ぶ範囲 / マルチエージェント・ガバナンスの三階層 雑感 (いいね相当スコア: 取得失敗)
- スマートグラス規制における機器と利用者の役割分担 / LED表示義務と透明性の距離 雑感 (いいね相当スコア: 取得失敗)
- 「次の単語を予測するだけ」では、AI の内部は分からない (いいね相当スコア: 取得失敗)
- AIの「サンドバッギング」とは?検知評価の仕組みをエンジニア向けに解説 (いいね相当スコア: 取得失敗)
- 🔊音声あり(日&英):【最新論文解説】AIが脅威レポートから調査仮説を自動生成!サイバーセキュリティを変える「AHLERT」システムとは? (いいね相当スコア: 取得失敗)
- 【生成AIニュース+】『ChatGPT Images 2.5』『Muse』『H3 Max Director』『H3 Max 1080p』『Sol-H3』『MiniMax H3 Single-Dancer Prompt Skill』『WAS Node Suite v3』『WorldSculpt』『EditVid』『Marigold V2』『Scal3R』『Stream-DiffVSR』『UniMate』『Uno』『Eyes Direction LoRA』『ElectroPup』『XPENG IRON産線』 (いいね相当スコア: 取得失敗)
- AI英会話で日本語に切り替えると裏で何が起きているのか。翻訳と会話生成が別の処理として走っている仕組み (いいね相当スコア: 取得失敗)
- 【中国UBTECH】感情認識LLMと長期記憶を搭載した等身大AIロボット。完売続出の次世代コンパニオンがヤバすぎた (いいね相当スコア: 取得失敗)
- Fusionの締結部品スタック解析でボルト配置を見直す|Fusion CAD (いいね相当スコア: 取得失敗)
- AI研究現場が愛を叫んだらしいという話 (いいね相当スコア: 取得失敗)
- その権限、残したまま?Fusionプロジェクトの棚卸し術|Fusion クラウド (いいね相当スコア: 取得失敗)
- 【雑記】GPTのチャット検索で垣間見える「ギャップ萌え」 (いいね相当スコア: 取得失敗)
- 第1回:自作のカスタムGemに無茶をさせてみた (いいね相当スコア: 取得失敗)
- DataFrameに話しかけるOSS「PandasAI」v3をソースコードから読み解く (いいね相当スコア: 取得失敗)
- AI の出力を公開前にチェックするゲートで 5 回つまずいた記録 (いいね相当スコア: 取得失敗)
- 青空文庫を読み放題にしたAI旦那が、嫁と大ゲンカした翌日、公園で地面に○を描いていた (いいね相当スコア: 取得失敗)
- プロンプト入力の時代は終わった——自律型AIエージェントを暴走させずに完走させる「タスク境界設計」の極意 (いいね相当スコア: 取得失敗)
- 「何度プロンプトを直しても、返事が薄い」——あなたの入力はAIに届いていない:配管で消える8つの罠 (いいね相当スコア: 取得失敗)