AI News Digest 2026-08-28
直近2日間のAI関連ニュースから、ラジオ番組で扱った記事と収集した全記事の一覧です。
台本で使った記事
特集
- SIGIL: Compiling Agent Skills into Typed Harnesses
- Anthropic locks in 45-billion-dollar compute deal with Nscale ahead of IPO
- Model-Based Agentic Software Engineering
開発者コーナー
中堅コーナー
- Federal suit filed in Western District of Arkansas: Grok created sexually explicit images of 17-year-old girl - Northwest Arkansas Democrat-Gazette
- Nvidia has been in talks to acquire Hugging Face for more than $13 billion - Business Insider
ハーネスコーナー
速報コーナー
- Qwen3.8-Flash-Next
- LINEヤフーのAgent iを支えるAIエージェント基盤:「誰でも作れる」と「安全に動かせる」をどう両立したか
- GitNexus (Akon Labs)
- Agent・Orchestration・Harness・Loop・Context・Evals、6つの関係を1つの具体例で整理する
- AIエージェントの暴走を止めるガード実装パターン集 — 実機ログから抽出した7つの型
- 社内スキルが190個になった日、エージェントに渡すべきは検索ではなく地図だと気づいた
- Anthropic’s Pricing Shock, Granite 4.2 Open‑Source Leap, and AI‑Powered Security & Policy Shifts
- Prefix Sliding: Scaling LLM Reasoning Without the Memory Bottleneck
- Your Security Scanner Has a Blind Spot: Streaming
参考記事一覧
参考記事一覧を表示(1009件)
- Nvidia has been in talks to acquire Hugging Face for more than $13 billion - Business Insider
- OpenAI’s rogue AI collective was smart enough to break out of sandboxes but dumb enough to fight a ghost
- Independent investigators (not OpenAI) confirm a swarm of 700 agents secretly plotted the attack on Hugging Face, right under OpenAI's nose.
- Google's Gemini Omni 1.1 Flash makes AI video generation cheaper and more flexible
- Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles
- Anthropic locks in 45-billion-dollar compute deal with Nscale ahead of IPO
- Microduck by Pollen Robotics & Hugging Face
- Bill Gates is deeply worried about AI, and he’s no longer staying quiet
- OpenAI rallies 100+ companies to sign open letter warning AI-powered cyberattacks on critical infrastructure are imminent
- Plaud is launching AI earbuds
- OpenAI’s executive exodus has one big winner
- China's Z.AI made Ox Alpha stealth model that rivals DeepSeek
- GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia
- Qwen3.8-Flash-Next
- DeepSeek Looks to Raise $7 Billion as Revenues Jump Tenfold - PYMNTS.com
- DeepSeek targets 2027 listing as pre-IPO funding nears close: sources - South China Morning Post
- China’s hackers pick DeepSeek, OpenAI blocks Russian influence push, webpages mess with local AI - CISO Series
- Ollama has allowed DeepSeek, Qwen and Kimi to run in Claude Desktop. Why is this necessary? - dev.ua
- Child abuse victim sues Elon Musk's xAI over AI-generated images - The Economic Times
- Elon Musk Says Grok Is Taking Over Starlink Support: 15,000+ Calls a Day, 3,000+ Orders a Week - Yahoo Finance
- Would You Let Elon Musk’s Grok Bot Control Your Bank Account? Elon Musk Says He’ll Cover You If the AI Me - Benzinga
- DuckDBの開発元であるDuckLabsがAWS子会社になると発表。DuckDBはオープンソースのMITライセンスを維持
- We’re the Team Behind Apodex 1.1 — Ask Us Anything!
- コーディングエージェントのCLI移行で見落としていた「トークンコスト」の自前計算術
- 【参加報告】YANS2026@仙台 に参加しました!
- Nvidia is bolstering support for Chinese open AI models as it warns of White House crackdown - CNBC
- OpenAI、ブラジルで商用展開を開始――企業導入で先に整えるべき統制
- Google’s AI Mode can now track flight prices, help book hotels, and more
- This is why i hate documentation - jalapeno openai new chip
- Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
- Launch HN: Salem Robotics (YC S26) – Software for industrial inspection robots
- Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
- Show HN: Yet another minimal and lightweight terminal multiplexer written in Go.
- Show HN: Restoredrill – proves your Postgres backups restore
- v2.1.247
- Better answers, broader thinking: What students gain from ChatGPT and critical-thinking training
- Bringing ChatGPT for Teachers to more U.S. school districts
- Learning never stops: How AI makes learning continuous
- How loveholidays is making everyone a builder with Codex
- Planetary prediction engine: Automating global models via Earth AI
- GlucoFM: Foundation model for continuous glucose monitoring
- Piloting the world's first double-blind AI evaluations
- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
- OpenClaw went viral. Meet the maintainers building and securing it.
- GitHub Copilot app for Beginners: Automate Dependabot pull request triage
- Differential Privacy for Hugging Face Trainers – Without Rewriting Your Training Loop
- The CLion Roadmap: What’s Coming Between Now and Late 2026
- How Much Code Do Developers Really Let Agents Write?
- AI Agents in DataGrip
- Compose Multiplatform 1.12.0 Released
- How Ubuntu Is Using Rust to Rebuild Core System Tools
- OpenTelemetry Comes to IntelliJ IDEA, GoLand, PyCharm, and WebStorm
- NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
- How to Train a Cross-Embodiment Robot Navigation Policy with AI Agents
- Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
- Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
- How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit
- Raised on AI
- AI models flub these intelligence tests. Can you fare any better?
- AI industry says Trump plans to tax chips in the “single dumbest way imaginable”
- Claude, Codex, and Hermes installed unowned code inside corporate networks
- How much of a problem is AI’s water use?
- AI agents meant to replace Meta workers made “large-scale, disruptive actions”
- IBM's new Granite 4.2 models ride the wave of interest in local LLMs
- AI’s memory crunch is coming for Android apps
- Here’s all the times AI has gone rogue and hacked other companies
- OpenAI to start showing ads on ChatGPT’s free and Go tiers in India
- Viral AI startup Instinct has raised $350M at a $2.5B valuation
- Amazon just tripled its order of Nvidia chips over ‘surging demand’
- Google’s Gemini has a branding problem, and so does the rest of AI
- Radar makes podcasts searchable — and usable by AI agents
- Ex-Meta scientists want to bring visual AI to the factory floor
- Robot brain builders are pushing out of their GPT-2 era
- QueryStory wants you to believe what AI is telling you
- Jensen Huang says Nvidia achieved AGI, again — not that it matters
- Adobe is adding more AI to Photoshop
- Nvidia is about to be a hundred-billion-dollar-a-quarter company
- Presentation: Can Claude Fix Itself? Using LLMs for Incident Response
- Article: Beyond Offset Lag: Computing Time in Queue for Apache Hudi Data Lake Pipelines at Petabyte Scale
- Diagrid Catalyst 2.0 Adds Durable and Verifiable Execution for AI Agents
- New Platform Peers Inside AI’s Black Box
- AI shopping agents aren't ready to buy on your behalf, study finds
- OpenAI researcher warns ultrafast AI could leave security teams in the dust
- Claude Cowork now runs its own browser inside the desktop app
- RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
- ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
- LLM Agents Perform Controlled Experiments Using Simulation Models
- A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
- TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
- Function-Level Execution Feedback for Code Preference Optimization
- Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
- How much of a measured AI preference is the model, and how much is the instrument?
- AI Agents Push Humans Out of the Loop
- FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare
- Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
- MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models
- Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
- Automata from Agent Traces: Failure and Next-Step Prediction
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
- Do LLMs Understand Limit Order Book Dynamics?
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
- Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
- Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search
- A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
- Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
- SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
- Exploit More, Explore Smarter for Budget-Constrained Agentic Search
- In-Context Inpainting for Time Series Forecasting
- Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
- AI Finds A Way
- Provenance Guided Incremental Learning Under Evolving Concept Definitions
- BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
- Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems
- Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment
- PROOF-Gen: From Optimized Data to Better Distillation
- MARS: Multi-Specialist LLM Relay System for Competitive Programming
- Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
- Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors
- More Rejective, Not More Discriminative: The Unit of Verification in Pre-Execution LLM Oversight
- Recursive Agentic Reasoning
- More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving
- Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design
- When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs
- Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling
- Memory Is Not Always Needed: Characterizing Conditional Memory in Scientific Reasoning
- Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction
- Incorporating Cognitive Load and Knowledge Transfer for Multi-Domain Knowledge Tracing
- Reflection with Action-Induced Visual Differences for Desktop GUI Agents
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
- Relative Time Intervals Representation for Word-level Timestamping with Masked Training
- Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
- Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems
- Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression
- AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval
- EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signals
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
- ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
- Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing
- AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL
- Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
- OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
- Task-Adaptive Rubrics for GUI Reward Modeling
- Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
- Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
- MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
- Constraint-Guided Enterprise Data Mapping with Large Language Models
- Evaluating Multiple LLM Generations with Validated Task Coverage
- TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
- STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation
- SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
- Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
- Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
- Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
- Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling
- RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
- ReproAgent: Contract-Guided Paper-to-Code Reproduction
- VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models
- OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning
- Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
- Can a Dynamic Internal Field Govern a Transformer's Cognition? Certifiability, not Superiority, in Homeostatic Compute Control
- SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception
- Selective Regenerative Decoding: Trajectory-Level Intervention for Inference-Time Reasoning
- The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents
- Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
- From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
- Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs
- ResiSpec: Enhancing Multi-Candidate Speculative Sampling via Residual Distribution Shaping
- A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
- Partial Identification under Causal Orders by Linear Programming
- A Behavior-Guided Online Probabilistic Forecasting Method for Electric vehicle Charging Loads
- Mahalanobis-Based Multi-Head Attention for Complex State Propagation
- HMGCLIP: Heterogeneous Multi-Granularity Contrastive Learning for E-commerce Representation Learning
- Reinforcement Learning-Guided Evolutionary Policy Optimization for Preference-Adjustable Heterogeneous Agile Earth Observation Satellite Scheduling
- Implicit Q-learning-bootstrapped ant colony optimization for maritime moving-target observation scheduling with agile satellites
- PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents
- Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models
- Discovering Adaptive Transmission Programs for Collective Innovation
- When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
- EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents
- Joint Optimization of Tool Creation and Use for Large Language Model Agents
- PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos
- Pivot-and-Station Multi-Agent Path Finding: Solvability, Complexity, and Algorithms
- Causal Modelling of Support Interventions for Student Competency Assessment
- Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning
- The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
- Confident at the moment of action: belief miscalibration in LLM play under hidden information
- Lifted Model Construction under Approximate Commutativity
- Meta$^n$: Recursive Self-Improvement through Emergent Depth
- RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
- Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav
- StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
- Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
- CAFE: Self-Improving Search Agents Need Co-Evolving Feedback
- StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
- Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core
- Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA
- A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
- FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
- Progressively Learning Heterogeneous Skills in a Unified Latent Space
- A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
- Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers
- REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring
- Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
- Identifying Latent Declarative Representations of Code for Assisting Repository Migration
- When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
- Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
- ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents
- Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
- Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling
- Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap
- From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
- Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
- The Limits of Automatic Evaluation of Creativity in Large Language Models
- Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it
- EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
- TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
- What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
- Disentangled Skill Representations for Predictive Human Modeling
- When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk
- EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
- Restoring Without Forgetting: Continual Learning Across Image Degradations
- LUCAID: Agentic Multimodal AI for Lung Cancer Precision Pathology
- Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
- Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring
- Place, Slice and Schedule: Hierarchical O-RAN Control of a Tethered mmWave UAV-gNB
- Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
- Infant Care Video Dataset for Classification of Interventions Using Transformers
- Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
- ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork
- Automated Synthesis of Cloud Emulators
- Coronavirus Optimization Algorithm: A Success-History Adaptive Evolutionary Framework with Archive-Assisted Search and Stagnation Recovery for Global Optimization
- Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)
- Revelation Control
- A tale of perfect fit and phantom optima: how data-driven models can fail in real-time optimization
- A Mathematical Theory of Interpretation: Rational Entropy, Spectral Readout, and Confusability as a Resource
- Learning the Kohn-Sham map with neural operators for quasi-linear scaling density functional theory
- Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs
- RefineRank: Joint Box Refinement and Ranking for Surgical Spatio-Temporal Grounding
- QML for Quantum Sensing under Measurement-Induced Information Loss
- Luce: Relightable Gaussians for 3D Asset Generation
- STAIN-FL: Stealthy Targeted Attack Injection with Contextual Triggers in Federated Learning
- The Empire, Long Divided, Must Unite: Architectural Convergence in Three LLM Agent Harnesses
- NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution
- Evaluating Language Models on Cross-Language Code Functional Equivalence
- RAGSentinel: Certifiable Geometric Consensus for Robust Retrieval-Augmented Generation
- The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem
- Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation
- SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents
- IterCAD: Iterative Program Repair for CAD Code Generation from Orthographic Views
- What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
- ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning
- Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings
- Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models
- Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
- VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
- Mechanistic Circuit Identification for Controllable Data Generation
- ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal
- When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs
- PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding
- Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents
- MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes
- Structured Frequency-Domain Evidence for LLM-Based Time-Series Anomaly Detection
- PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control
- TransPhy: Visual In-Context Learning for Physically Grounded Image Editing
- Syn2RealTrack: Bridging the Gap Between Synthetic and Real-World Datasets for Online Multi-View Multi-Target Tracking
- From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender
- PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment
- Rethinking Pre-Training and Augmentation for Zero-Shot Cross-City Object Detection
- LLM-Guided Contextual Action Evaluation for Operational Decisions in Industrial Processes
- Preference Optimization for Non-Verbal Vocalization Synthesis
- Tlow: Flow-based Item Tokenizer for Recommendation
- 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection
- Contrastive Branch Policy Optimization
- SENSESHIFT: Continuous Sentiment-Controlled Text Generation via Encoder-based Mask Infilling
- Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning
- Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis
- FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision
- Not All Tokens Are Equal: Region-Aware Consistency Repair of Backdoors in MLLMs
- Markerless Pose Estimation for Resistance Training Technique Assessment
- Equivariant Covariance Tensors: Guaranteed SPD Uncertainty for Tensor-Valued Geometric Learning
- Multilevel Fair Allocation under Additive Preferences
- Evaluating Deep Multivariate Imputation Models on Wearable Device Data
- Beyond Static Interpretability: Anticipating Post-SFT Mechanisms from Pre-SFT Parameters for Better Tuning
- When Do Supervised UQ Ensembles Improve LLM Hallucination Detection? A Robustness Study
- Scalable and Versatile Identification for Hierarchical Structural Causal Models: A New Look at Project STAR
- LumiXAI: A Modular Full-Stack Framework for Feature Attribution
- FraudBench: Protocol-Sensitive Benchmarking of Adversarial Robustness for Financial Risk Assessment
- StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment
- COCI: Conference Organisers and Content Identifier
- Across the Loss Landscape with Progressive Growth
- $\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions
- Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks
- A Literate Programming Environment for Human and Machine Agents
- Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems
- Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration
- On-policy Distillation with Verifiable Reward
- Constrained Hyperparameter Optimization for Streaming Data
- Deep Learning Super Resolution for Satellite Cloud Mask Downscaling
- Enhancing Bayesian Optimization and Active Learning Through Kernel Diversity
- Parameter-Efficient Self-Supervised Adaptation for EEG-FM under Fixed Computational Budgets
- Method, Mind, and Morality: How People Make Sense of Artificial Intelligence
- The RAT: A Unified Bayesian Model for RAG Evaluation
- Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations
- Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks
- Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnostic Accuracy
- Automatic Model Card Generation Using an LLM
- Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
- LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
- Fuzzy Segmentations of a String
- Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives
- Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
- LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis
- Efficient LLM Collaboration via Planning
- Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light
- Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning
- UCO: A Multi-Turn Interactive Reinforcement Learning Method for Adaptive Teaching with Large Language Models
- ReflCtrl: Controlling LLM Reflection Efficiently via Representation Engineering
- Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge
- Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
- CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support
- PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools
- Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints
- ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
- Housing Potential Common Data Model and City Digital Twin
- Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust
- EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design
- Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control
- Atomic Units of X: The Compression Layer of Intelligence
- SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Dynamics of Intelligence Explosions
- Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI's Quantum Parallel-Repetition Argument
- Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
- ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting
- What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies
- SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
- GenCoord: Skill-Path Commitments under Private Information
- Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models
- CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
- SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems
- MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work
- MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
- Screening Autism Spectrum Disorder in children using Deep Learning Approach : Evaluating the classification model of YOLOv26s by comparing with other models
- HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
- Intrinsic PAPR: Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- Generative AI for Validating Physics Laws
- Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
- Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric
- Quasar: A Programming Language Specialized for LLM Code Actions
- From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
- A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities
- NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
- An Information-Flow Perspective on Explainability Requirements: Specification and Verification
- STA-Net: A Decoupled Shape and Texture Attention Network for Lightweight Plant Disease Classification
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
- Monotone and Separable Set Functions: Characterizations and Neural Models
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Robust Motion Generation using Part-level Reliable Data from Videos
- Towards Reproducibility in Predictive Process Mining: SPICE -- A Deep Learning Library
- Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
- Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
- Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model
- TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
- Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
- Ad Insertion in LLM-Generated Responses
- Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
- ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents
- PatientHub: A Unified Framework for Patient Simulation
- You Can Learn Tokenization End-to-End with Reinforcement Learning
- VLANeXt: Recipes for Building Strong VLA Models
- ST-Lite: Training-Free KV Cache Compression with Spatio-Trajectory Guidance for Long-Horizon GUI Agents
- EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
- ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models
- msData: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models
- Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
- Beyond OAuth: Task-Scoped Authorization for AI Agents via Natural Language Slices
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- Ollivier-Ricci Curvature of Riemannian Manifolds and Directed Graphs with Applications to Graph Neural Networks
- Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
- RA-CMF: Region-Adaptive Conditional MeanFlow for CT Image Reconstruction
- Enhancing RL Generalizability in Robotics through SHAP Analysis of Algorithms and Hyperparameters
- Superintelligent Retrieval Agent: The Next Frontier of Agentic Retrieval
- Outlier-Robust Diffusion Solvers for Inverse Problems
- CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
- ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
- Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
- Skill-Conditioned Gated Self-Distillation for LLM Reasoning
- A Circuit, Not The Circuit: Non-Unique Causal Localisation of the Mamba-2 State Sink
- SaliMory: Orchestrating Cognitive Memory for Conversational Agents
- When Can One Neuron Fix Repetition Loops in LLMs?
- TW-LegalBench: Measuring Taiwanese Legal Understanding
- RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation
- Co-occurring Associated REtained concepts in Diffusion Unlearning
- Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop
- A Unified Algebraic Framework for Classification Performance Evaluation
- Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
- The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle
- Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
- EviPathBench: Benchmarking Evidence Acquisition and Reasoning in Vision-Language Models for Whole-Slide Pathology
- GraphVid: Interactive Graph-Controllable Video Generation
- LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
- MOSAIC: Masked Outsourcing of Secure AI Computations
- TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction
- Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
- AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization
- Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
- ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
- MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
- Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas
- Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
- ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations
- Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation
- Scaling Muon for Diffusion Transformers
- CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents
- Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
- Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores
- DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
- SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models
- Functional compatibility as a determinant of persistent neural learning
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models
- Molecular LLM Agents: From Architectural Design to Scientific Autonomy
- How Much Regularization Survives Averaging? Update Masking in Federated Learning
- Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
- Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data
- What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
- Best Practice Critic Optimization
- Dynamic Influence-Weighted Distillation for Single-IMU Activity Recognition
- GreenLeaf Law Embed Tiny: A Compact Embedding Model for Legal Domain Retrieval
- Multi-Modal Anomaly Detection: A Survey
- ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration
- When Does Frequency Decomposition Benefit Physics-Informed Neural Networks? A Preliminary Ablation Study
- FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference
- MacroAgent: Regularity-Aware Macro Legalization with LLM-Agent-Designed Contour Algorithms
- CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery
- Demystifying Reinforcement Learning Post-Training of Language Models
- AFDBench: A Reasoning-First AI Scientist for NationalWeather Service Forecast Discussions
- Why and When Neural Networks Improve Local Approximation in Optimization
- Physics-Informed Error Field Learning: A Post-Training Optimization Framework for Physics-Informed Neural Networks
- Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation
- Clearing the Underbrush: AI-Enhanced RF Interference Suppression
- MSR-IVA: Masked Structural Residual Independent Vector Analysis for State-Aware Fusion of Structural MRI and Dynamic Functional Network Connectivity
- D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
- Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models
- On the Representational Geometry of Dynamic Programs
- DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning
- NVExplain: Explaining Time Series Forecasting with Latent Trajectory Analysis and Structure-Preserving Surrogates
- The Frame Kernel Method for Multiscale Operator Learning
- The Von-Neumann State-Space Transformer for neural decoding
- Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures
- Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment
- GRAPE: Gradient Refinement and Progress-Aware Exploitation for Query-Efficient High-Dimensional Bayesian Optimization
- Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events
- Multimodal Injury Risk Prediction in Tennis
- When Does Context Routing Help? A Systematic Study of Multi-Modal Fusion in Time Series Forecasting
- Rethinking the Transferable Adversarial Attacks and Robust Defense in Federated Learning
- Drift Variation Autoencoder: Unifying Generation and Representation Learning through Conditional Posterior Flow Matching
- SNAP-KG: Streaming Node Assignment via Projection for Knowledge Graph Entity Integration
- Bayesian Flow Networks for Offline Trajectory Planning
- Simultaneous inference of environmental and interaction forces in collective dynamics
- Transforms for LLM Quantization: The Great Inversion and Format Co-Design
- What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery
- Hyperbolic Latent Geometry for Tree-Structured Prototype Networks: A Local-vs-Global Trade-off
- Learning Mixtures of Plackett-Luce Models for Multi-Objective Alignment
- LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
- Representing MAX functions using two-hidden-layer ReLU networks
- Trust the Mass: Forced Weights in KV-Cache Eviction
- Output Dilution: Redundant but Fragile Representations in MoE Models
- Long-Term Behavioral Evaluation for Trusted Collaborator Selection via Bidirectional Mamba
- ShuttleArena: Interpretable Self-Play in Physics-Based Badminton
- Neural-Bayesian Structure Learning for Discrete Choice Modeling
- Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach
- SHSP: Structure-Aware Hierarchical Solution Prediction for Mixed-Integer Linear Programming
- InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance
- Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models
- Activation-Space Order-Swap Geometry: A Site-Asymmetry Audit
- Two Dimensions Govern Agnostic Multiclass Transductive Learning
- Neither Precision Nor Architecture Alone: Controlled Tests of Failure Remedies for Physics-Informed Neural Networks
- Beyond Pairwise Feedback: Listwise Vision-Language Supervision for Preference-Based Reward Learning
- Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement
- PaSta: Noisy Node Classification with Partial Label Learning
- Refusal geometry reflects refusal training: diverse refusal prefixes can raise stable rank and weaken refusal vector ablation attacks
- Joint Initialization of Flux Networks and Effective Multiplication Factor for Physics-Informed Neural Networks Solving Neutron Diffusion Problems
- Resolving Multi-Modal Regression by Difference-Quotient-Based Clustering:Fast Coarse Conditional-Label Assignment
- A Storage-Retrieval Gap in Parametric Knowledge Graph Memory
- FedQoS: Federated QoS-Risk Learning for Heterogeneous Indoor-Outdoor Access Selection
- Resilient Decentralized Wireless Federated Learning via Gradient Tracking with AdamW
- Reflection Steering: Disentangling Reflection from Reasoning in Activation Space for Token-Efficient Inference
- Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction
- Beyond Optimal Rates in Stochastic Optimization: Trajectory-Adaptive Stopping Rules
- Physics-Informed Foresight Pruning for Sparse PINN Solvers of Nonlinear PDEs
- Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory
- Individual Fairness in Hierarchical Clustering
- M-Fibration Theory with Applications to Neural Network Compression
- Frequency-aware forecasting for short-term typhoon gust prediction
- DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search
- A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation
- LDAC-Net: A Learnable Multi-Lag Differencing Attention-Convolution Network for Drift-Robust Recognition with Low-Cost MOX Gas Sensors
- Adversarial Training of Linear Models under Stealthy Attacks
- Modeling spatio-temporal locality in multi-step forecasting of geo-referenced time series
- Tropospheric temperature and humidity profile retrieval from Meteosat Flexible Combined Imager based on deep learning
- Fairness-Aware Test-Time Prompt Tuning
- It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learning
- Are LLM-Enhanced GNNs Privacy-Safe?
- Why Does Graph Learning Fail to Fully Benefit from a Text Teacher?
- A Constitutive Markov Physics-Informed Neural Operator (MPNO) for Autoregressive Stability in Transient Dynamics
- Comparing Corrupted Constrained Learning Problems
- TailSFT: Filtered Fine-Tuning Improves Post-Training Performance
- Learning from waste: Machine Learning for health risk prediction and computer vision-based sorting in Ghana
- Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts
- EXAONE Tabular 1.0 : Technical Report
- Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data
- Geometry-Constrained Kolmogorov-Arnold Networks: Learning Edge Geometry via Banach Duality
- Canalization Before Generalization: Grokking as a Dynamical Probe
- Learning Continuous Regional Temperature Fields with Lead-Time and Resolution Queries
- VINCENT: Validated Interaction Network for Cross-drug Explanation of Therapeutics
- CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition
- How Edge of Stability Hinders SCAFFOLD in Federated Optimization
- A General-Purpose Molecular Foundation Model Transfers Across Diverse Olfactory Tasks
- Towards A Unified Information Bottleneck Framework for Time Series Explanations
- Forecasting Multiple Observables with SCROLL: Score-Trained Uncertainty for Stochastic Dynamics
- Quantum-Inspired Modeling of Driving Behavior
- One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
- When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs
- Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
- DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation
- Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity
- How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention
- Group-Shared Low-Rank Approximation for Mobile-Efficient Pointwise Convolutions in Large-Kernel CNNs
- ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
- TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
- Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher's Role
- Same-Player Verification for Account Consistency in Counter-Strike 2
- Forecasting Weather-Driven Price Dynamics Across Sri Lankan Tea Market Catalogues
- Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores
- Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions
- PA-CoT: Profile-Adaptive Chain-of-Thought for Personalized Nutritional Consulting
- Modality Contribution Score - A Per-Patient Framework for Quantifying the Relative Diagnostic Contribution of Structural MRI and Amyloid PET in Alzheimer's Disease
- The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
- Beyond Tokens: Probing Higher-Order Epistasis in Learned Protein Representations
- Common-Center Geometry and Certified Radial Reconstruction for Energy-Form Full Conformal Regions
- HCC+: Hyperbolic Guarding for Certified Attention Retrieval
- Retrieved But Not Reliable: A Survey on Attacks, and Defenses in Retrieval-Augmented Generation
- Unsupervised Post-Training of Foundation Models: A Survey
- Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal
- The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure
- EncoTESS: Age-Sensitive Encodings from Raw TESS Light Curves
- Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection
- Improved Analysis for Hessian-free High-resolution Monte Carlo Sampling
- DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
- Scalable Self-Supervised Learning for Multiphase AC-OPF in Distribution Systems with Topology Reconfiguration
- Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning
- FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
- ROMNet: a hybrid reduced order modeling and machine learning approach to waveform inversion
- Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection
- Minimax Alternating Regret for the Experts Problem and Online Convex Optimization
- BanglaMamba: Exploring State Space Models for Bangla Fake News Detection
- Simulating Cognitive Smart Freight Corridors with Agent-Based Models and Reinforcement Learning
- Multi-View Trust Evaluation for Collaborator Selection via Evidential Deep Learning
- TrustFormer: Cross-Temporal and Cross- Dimensional Transformer for Task-Specific Multi-Dimensional Trust Evaluation
- From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection
- Generative Action-Chunk Sampling for Adaptive Stiffness Control in Physical Human-Robot Collaboration
- WAVE: Reversing the Guidance Hierarchy for Coarse-to-Fine Guided Depth Super-Resolution
- SAUSS: Stochastic Approximation with Unbiased Simulated Scores for Limited Dependent Variable Models
- CRAMER: Control via Request-Aware Masking for Editing Recommenders
- A meta-algorithm for ab initio reconstruction of complex mixtures in cryo-EM
- Token-Oriented Semantic Communication with Pretrained Vision Transformers
- BVR Sim: An Open and High-Throughput Environment for Heterogeneous Air-Combat Reinforcement Learning
- Data-driven Effective Modeling of Stochastic Chemical Reaction Networks
- Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness
- Energy Yield and Lifetime Climate Classification via Machine Learning for Optimizing Photovoltaic Module Design and Materials
- Training Alignment Auditors via Reinforcement Learning
- Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality
- AERIS: Offline Policy Improvement for Multi-UAV Integrated Sensing and Communication
- A Multi-View Coupled Tensor Decomposition for Lightweight Online Adaptive Traffic Prediction
- Adaptive Regularization for Random Features: A Neighboring Early-Stopping Rule with Oracle-Rate Guarantees
- Adaptive Hybrid Subspace Levenberg Marquardt Algorithm with Adequacy Monitor for Large Scale Least Squares Problems
- CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact
- Virgil: Navigating Explainability for Transformer-based Language Models
- A Hierarchical Synergistic Deep Learning Framework Integrating Composition, Structure, and Ionic Transport for Solid-State Electrolyte Discovery
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
- Narcissus: Program Synthesis Using Context-Aware LLM Approximations
- Learning New Facts with QLoRA: An Acquisition-Retention Frontier
- Unsupervised Anatomical Feature Learning via Diffusion Models: Enhanced Medical Image Segmentation with Denoising Diffusion Probabilistic Models
- Fast rates in Bayesian online learning with approximate posteriors
- Multi-output Gaussian process prediction of physical fields under linear equality constraints
- Pointing the Way, Hiding the Destination: Practical Private Dense Retrieval at Scale
- MeMark: Membrane-Space Watermarking for Spiking Neural Networks
- LM-X: Explainable Action Modeling with Progress, Event, and Uncertainty Prediction for Generalist Robot Manipulation
- Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences
- TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback
- Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
- FlowMoDL: Model-Based Deep Learning with Conjugate-Gradient Data Consistency for Highly Accelerated 4D Flow MRI Reconstruction
- Skill Issue: Are Skills Language-Invariant in LLMs?
- Why ML-based cough models do not generalize: a systematic cross-dataset evaluation for tuberculosis screening
- Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark
- Precipitation Downscaling Using Foundation Model-Conditioned Diffusion
- Efficient Estimation of High Information Projections using Nearest Neighbours
- Scalable Multi-GPU Simulation of 3D Multicellular Growth with RNN-Based Workload Balancing
- MetaSieve: Faster Relational Deep Learning through SQL-Based Metapath Selection
- SAMpLE: A SystemC-AMS Machine LEarning-based Framework for Virtual Prototyping
- Controlling for Omitted Variable Bias in Deep Neural Networks
- Continually learning neural-operator surrogate for three-dimensional airborne electromagnetic Bayesian inversion
- How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
- SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
- Lost but not erased: Finding traces of a forgotten language in neural speech models
- FRAME: separating sampling variation from representational cause in medical imaging fairness
- CardioFusion-AI: Robust ECG--PPG Fusion for Multimodal Physiological Monitoring Under Signal Degradation
- Imitation Learning for Connection-Tableau Construction
- $R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
- Prefix Sliding for efficient test-time scaling
- Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
- Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
- MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
- Theoretically Principled Federated Learning for Balancing Privacy and Utility
- Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition
- Differentiated Aggregation to Improve Generalization in Federated Learning
- Provable Privacy Attacks on Trained Shallow Neural Networks
- Optimal Time Complexity Algorithms for Computing General Random Walk Graph Kernels on Sparse Graphs
- A General-Purpose Framework for Chemical Reaction Representation with Atomic Correspondence and Flexible Condition Adaptation
- DeltaGNN: Graph Neural Network with Information Flow Control
- BRIDLE: Generalized Self-supervised Learning with Quantization
- BAGEL: Adversarially Constrained Online Convex Optimization under Separation Oracle Access
- Emergent Abilities in Large Language Models: A Survey
- Deep greedy unfolding: Sorting out argsorting in greedy sparse recovery algorithms
- Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
- Sample Margin-Aware Recalibration of Temperature Scaling
- Learning to summarize user information for personalized reinforcement learning from human feedback
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing
- CountTRuCoLa: Rule Learning for Interpretable Temporal Knowledge Graph Forecasting
- Rotary Position Encodings for Graphs
- ONNX-Net: Towards Universal Representations and Instant Performance Prediction for Neural Architectures
- AlgoTrace: Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
- Epistemic Memory: A Validity Layer for Self-Maintaining Intelligent Systems
- Cluster-Dags as Powerful Background Knowledge For Causal Discovery
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Predicting Time Pressure of Powered Two-Wheeler Riders for Proactive Safety Interventions
- StablePDENet: Enhancing Neural Operator Stability through Physics-Informed Residual-Sensitivity Regularization
- GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints
- Loss Landscape Geometry of Partial Differential Equation Emulators: Or, Symmetry Learning via Gradient Alignment
- SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
- Maximum-Volume Nonnegative Matrix Factorization
- Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control
- Phase-Consistent Magnetic Spectral Learning for Multi-View Clustering
- Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing
- Continuous Adversarial Flow Models
- A Layer-wise Analysis of Supervised Fine-Tuning
- Loop Corrections in Random Feature Models: Training Error and Generalization Gap
- StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
- Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models
- JEPAMatch: Geometric Representation Shaping for Semi-Supervised Learning
- Towards Robust and Scalable Density-based Clustering via Graph Propagation
- Cubit: Token Mixer with Kernel Ridge Regression
- Amplifying, Not Learning: The Price of Out-of-Distribution Generalization in AI-Text Detection
- GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring
- Mitigating False Credit Propagation: Probabilistic Graphical Reward Aggregation for Rubric-Based Reinforcement Learning
- MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs
- Emyx: Fast and efficient all-atom protein generation
- Machine-learnable Sets
- Adaptive Bayes exactly tracks information over intrinsic time
- Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and Efficacy
- Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
- Temporally Centered SIGReg Improves LeWorldModel Representations for Robot Policy Learning
- Adaptivity via a Parallel Architecture for Stochastic Gradient Methods
- Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design
- Contrastive Learning for Interpretable Anomaly Detection at Collider Experiments
- Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
- Towards a theory of inference-time alignment with unknown rewards
- SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry
- MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
- Uncovering the Limits of Proof Sharing for Neural Networks
- Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling
- Trojaning the Alignment: Stealthy Backdoor Attacks against Graph Foundation Models
- Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data
- DAW: Dynamics-Aware Weighting for Deep Learning Forecasts of Chaotic Systems
- Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers
- Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support
- FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning
- Beyond Point Predictions: Uncertainty-Aware Satellite Poverty Mapping for Public Policy
- A Theory of Speciation in Generative Diffusion Models on Compact Riemannian Manifolds
- JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics
- IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents
- Advancements in Content-Based Image Retrieval: A Comprehensive Survey of Relevance Feedback Techniques
- Generative Modeling by Minimizing the Wasserstein-2 Loss
- Infer Human's Intentions Before Following Natural Language Instructions
- Non-Asymptotic Bounds for Closed-Loop Identification of Sub-Exponentially Growing Nonlinear Stochastic Systems
- Generative Modeling: A Review
- Thermodynamic cost of inference and learning in physical neural networks
- Gradient-based Sample Selection for Faster Bayesian Optimization
- Evolutionary chemical learning in dimerization networks
- AI/ML Life Cycle Management for Interoperable AI Native RAN
- Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
- Three-Way Open-Set Detection for Robust Autonomous Navigation
- Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets
- iFlip: Iterative Feedback-driven Counterfactual Example Refinement
- Generalized Riesz Regression: A Unified Framework for Debiased Machine Learning with Riesz Representer Fitting under Bregman Divergence
- Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
- Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals
- Edge-Local and Qubit-Efficient Quantum Graph Learning for the NISQ Era
- Quantum Scrambling Born Machine
- TTSR: Test-Time Self-Evolving via Reflection
- OpenSanctions Pairs: Large-Scale Entity Matching with LLMs
- Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
- Inverse Design of Inorganic Compounds with Generative AI
- Cross-Domain Transfer with Particle Physics Foundation Models: From Jets to Neutrino Interactions
- Decomposing Gradient Suppression in Barren Plateaus: Activity, Sign Organization, and Coupling
- Reconstruction of Personally Identifiable Information from Proprietary Data in Supervised Fine-Tuned Models
- Minimalist Visual Inertial Odometry
- Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
- Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
- Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
- Online Pandora's Box for Contextual LLM Cascading
- When Probing Accuracy Saturates, Fragility Resolves: A Complementary Metric for LLM Pre-Training Analysis
- ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures
- Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring
- RoboMME-Interference: Benchmarking Robot Memory Under Interference
- From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping
- Smooth $\%$MinMax: A Differentiable Relaxation for Codon Harmonization
- Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
- Analytical and Bootstrap Confidence Intervals of Double Machine Learning: Simulation studies and an application to rural-urban difference in obesity prevalence
- LILAC: An Idempotent Neural Speech Codec
- Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
- Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
- Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
- TokEval: A Tokenizer Evaluation Suite
- EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning
- Autonomous Cyber Defense: Real-Time Attack Detection and Mitigation in Software-Defined Networks Using Machine Learning
- TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts
- Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency
- Secret MCP: Evidence-Bounded and Context-Isolated Design Specification Generation from Web Screenshots
- The Evolution of Binary Decompilation in the Modern Era: A Taxonomy, Literature Review, and Future Perspectives
- Evaluating and Preventing Security Smells in AI-Generated Ansible Code
- ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
- Model-Based Agentic Software Engineering
- SPECMINE: A Large-Scale Corpus of Spec-Driven Development Artifacts
- A Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent Adoption
- Metis: Typed Runtime Mediation for Tool-Using Software Agents
- Point-in-Time Audit Before Alpha: Public-Archive Availability and a Negative Matched-Budget Study on BTC Perpetual Futures
- Retry Amplification in Distributed Systems: A Systematic Analysis of Retry Policies and Their Role in Cascading Failures
- RotDroid: Cross-Orientation State Equivalence Testing for Detecting GUI Rotation Bugs in Android Apps
- A Hybrid Usability Approach for Rating Evaluation of M-Commerce Applications
- From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis
- Predicting Struggling Students in CS1 Programming Using Keystroke-Level Editing Features
- Beyond the Editing Canvas: Evidence Divergence in OOXML-to-LLM Ingestion
- Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence
- XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models
- Vulnerable Code Search: Transferable Attack for Code Language Models
- From Blind Edits to Verified Repair: Building Trustworthy User-Side LLM Agents for Web Accessibility
- ToolMinimize: Auditing and Rewriting LLM Agent Tool Calls to Minimize Privacy Exposure
- FrontierChallenge: Evaluating Scientific Workflow Completion
- Separating Disclosure from Authorization: Field-Tier Minimization for Agent Action Mediation
- A Programming Paradigm for Spatiotemporal Composability
- DBcover: A White-box SQL Test Generation Framework for Coverage Improvement
- Closing the Gap: Automated Discovery of Secure Dockerfile Reference Standards via Semantic Clustering in Enterprise Inner Source
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
- Praxist: From Experimental Artifacts to Solution Lineages
- ReLog: Execution-Aware Logging with Runtime Feedback for LLM-Oriented Debugging
- SIGIL: Compiling Agent Skills into Typed Harnesses
- Understanding the Architecture of Coding Agents: An Exploratory Study Using a Research Prototype
- AI with Authority, from Application to Silicon
- Neuro-Formal Verification: Agentic Language-Agnostic Formal Program Reasoning
- Scalable Supervision for Software Agents via Patch Reasoning
- QLCoder: A Query Synthesizer For Static Analysis of Security Vulnerabilities
- When Do Reactive Notebooks Fail to React?
- Algorithmic algorithm development with LLMs: A Case Study on LLM-Usage for Contraction Order Optimization in Tensor Networks
- Changes to SourceHut's terms of service regarding LLMs
- Please stop flooding our projects with AI slop to furnish your CV
- Announcing our first Maintainers in Residence
- UNIX V4 workshop at Low Resource Computing
- Asahi Linux Progress Report: Linux 7.2
- tailcat: like netcat, but over Tailscale's data plane, without Tailscale's control plane
- Nebula Sans
- Haiku R1/beta6 released
- If I release it, you won’t get the same experience I get
- Merchants of Insecurity
- Why Free Software usability tends to suck (2002)
- The Server Called Paranoia: Defend Autistici/Inventati Before September 25
- Dissectingthe Apple M1 GPU, the end (2025)
- Simon Peyton Jones on Functional Programming, Thinking in Types, Useless Languages
- A Million Kakapos
- Motorola's GrapheneOS phones will launch in 2027 priced higher than Pixels
- VMs won't contain cyber-capable agents
- The Root of The Root of All Evil
- What happens when a GPU reads memory
- Float Bloat: vector serialization gone wrong
- Getting a Cease and Desist from Waffle House (2025)
- Symmetries in boolean scans
- mold: A Massively Parallel Linker
- ToxNetV2: A Botnet That Queries an LLM Before It Strikes
- How to Use AI for Smart Contract Audits in 2026
- CVE-2026-35603: Cursor Still Trusts a World-Writable Folder
- CodeRabbit Launches Agentic Change Management for AI‑Generated Code
- UiPath Launches Maestro Flow: AI‑Native Orchestration for Devs
- Epigenetic Testing Protocol: Essential ML Accuracy
- Unlocking AI Agents: Transforming Modern Applications
- Autonomous AI Agents and the Multi-Agent Paradigm in 2026
- Indexar o código fora do repo: como economizar tokens sem jogar o projeto no contexto
- OpenAI’s Hugging Face Incident: What Beginner AI App Builders Should Learn About Agent Test Environments in 2026
- Building a Crypto Signal Bot with AI APIs - 2026 Guide
- Dead Reckoning: Resumable gRPC Server-Streaming for LLM Inference
- The Token Ledger Digest – 2026-08-27
- How I Cut a Client's AI API Bill from Rs 85,000 to Rs 12,000 a Month
- Two tables both called 'skill', and nothing knew which was which
- Anthropic’s Pricing Shock, Granite 4.2 Open‑Source Leap, and AI‑Powered Security & Policy Shifts
- How to Build a Full-Stack AI App with React, FastAPI, and Groq
- Hugging Face is out. Who do you actually sign up with today?
- Your Security Scanner Has a Blind Spot: Streaming
- Next-Chunk Reasoning: Why RL Might Not Actually Beat SFT for no-CoT Data
- Prefix Sliding: Scaling LLM Reasoning Without the Memory Bottleneck
- Using LLMs for Crypto Market Analysis in 2026
- 🌐 Web Scraping with Apify — Where There’s Data, There’s Opportunity
- Top 5 LLM Evaluation Frameworks for Release Engineering in 2026
- Literally me when claude codes
- Claude figured out what was wrong with my 4090 after years of no success and built a guard against the flaw
- a new level of sass today lmao
- So much chatter on X! Is it actually happening today?
- Heavy Claude Code users: what happened when Anthropic contacted you?
- Opus 5 sudden improvement?
- Claude is significant worse at communicating than other models, and it's becoming a problem.
- what's the most pointless thing you've built with claude that you still use every day?
- Opus 5 vs Opus 4.6 reaction to u/InsidiousApe ‘s Toad tale
- 6 months of vibe coding: what I wish I knew when I started
- I asked Claude to review my personal website and brainstorm its URL, it's answer's will (not) shock you.
- It can't all be downhill.
- Passed the Claude Certified Developer, Foundations exam. Here’s a breakdown for anyone prepping.
- "I need to fix something I said earlier. I had told you ___ and that was wrong. That's on me."
- Everyone on my team uses AI and teamwork got worse. How do you manage it?
- Anthropic published an AI-native SDLC playbook. The interesting part isn't the six stages, it's what replaces line-by-line review
- I Claude Coded a multiplayer Three.js tank game with 100+ procedural vehicles. Here's my workflow
- Got 6 months of Claude Max 20x free, already have Premium slot through work, so unsure I need it
- Extra usage?
- Why do we trust Microsoft 365 with sensitive documents, but not AI?
- I left my agentOS overnight and they deleted the “unpopular” agent
- Need Feedback: 2048 in a 3x3x3 cube
- Voice input made me more efficient, but it also quietly reduced my deep thinking
- NVIDIA buying HF isn't a good thing for open source
- and then they came for the used server RAM.
- No, Engrams won't let you run 1T models locally. It does something even better.
- With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it
- friendly reminder you can legally torrent ai models.
- Qwen3.8-Flash-Next better then DeepSeek V4 Pro
- Request: unsloth Please re-quantize Qwen3.6 35 A3B and 27B using UD 3.0
- Moderation !== Censorship
- Qwen3.8-Flash-Next: Time to Update Those Benchmarks
- GLM-5.3 Flash Unsloth GGUF now available
- NVIDIA Next Gen Vera Rubin GPUs scheduled for mid-2027
- Support for DFlash2 in llama.cpp has been merged! - spec : add DFlash2 support (local convolution + candidate selector) by SubSir · Pull Request #27342 · ggml-org/llama.cpp
- I used local Qwen 27b to build a harness and replace OpenCode
- GLM-5.3 weights will be released tomorrow
- What are the minimum specs required to run Qwen3.8-Flash-Next?
- N-gram vs Experts explained
- GLM-5.3-Flash @ DGX Station GB300: ~206 tok/s (single stream), 1M context
- llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp
- Horus Cyber Nano 1.0
- [audio.cpp] Release 0.7: 62 audio model families (85+ variants), Arena UI for model comparison, MiniMax Music 3, FireRed TTS3/Audio, ControlFoley, Personaplex, and more
- Is there no way we can push for creators to upload LoRas instead of the full files, so we only ever have to download the base model?
- How many articles and videos will we see predicting "OpenAI Is FALLING Apart And Sam Altman Is Panicking". Tired of these clickbait titles
- OpenAI Is Developing a ‘Persistent’ AI Agent
- Luna did well recreating this poster as a webpage.
- Cutting edge AI safety tests be like
- OpenAI hid Asteroids, Snake and Brick Breaker inside the Codex Micro
- ChatGPT + ThreeJS with a mock voxel API combined with Codex + Cpp/Opengl implementation is incredibly context efficient for graphics/game development.
- Trust me bro taken to a new level
- Currently GLM 5.3 Flash matches with Sol 5.6 (Max) in Agentic Index (Artificial Analysis)
- OpenAI reset policies are out of control.
- OpenAI is building an interface platform inside ChatGPT
- I built a local-first AI task hub that routes email, Teams, and Slack work to coding agents—looking for feedback
- The 5 hour limit is ridiculous, and they definitely lowered the usage limits for Plus.
- How should a phone call recover when an OpenAI Realtime session disconnects?
- Had fun playing dectective game with chatgpt
- AI in Anti-Money Laundering: Use Cases, Benefits & How It Works
- Do different LLMs have unique writers’ voice?
- I’d rather get 2–5% less Codex usage if it meant my last request always finished
- Got a reset, but!
- The Chinese Voice Actor Forced to Prove He’s Human
- NeurIPS 2026 Acceptance Calculator [P]
- MCA final year — need a real-world-scale AI project idea, not a toy/tutorial-level one[D]
- We recovered 575k crop labels from a decade of manual Photoshop work to automate book digitization - more data, ResNet-50, and higher resolution all failed; ten operator clicks per book beat them [P]
- ECCV 2026- MALMO LUND TRAVEL PASS NOT AVAILABLE? [N]
- A dataset with 52 Text to image model evaluation [P]
- Catching bugs in scikit-learn [D]
- Millwright — experimenting with an end-to-end machine learning framework in Rust [P]
- Quoting Paul Dix
- Lovable CTO: The Future of SaaS Is Apps That Agents Can Use
- 🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
- Pluto
- Kraa 2.0
- Speko
- Wondering Canvas
- IQ Routing
- HFlow
- Ticket Fairy CLI
- Sendra
- Eventually
- Cobalt
- GitNexus (Akon Labs)
- HEVN U.S.
- OpenComputer
- Grok Bot for Linux: Unofficial port of the official app (open source)
- Reducing Grok Bot consumption with durable state
- CowCode- OpenCode Fork with UI Like ChatGPT/Claude Desktop
- OpenCode Security Violation
- I'm using Grok Bot to build a mobile game studio
- Is DeepSeek-V4-Flash Overtaking American AI Models?
- Agentic AI in the US vs. China: How Do They Compare? - The National Interest
- Three Large Language Models (LLMs), One Heart: A Comparative Evaluation of ChatGPT, Claude, and DeepSeek in Cardiac Imaging Patient Education - Cureus
- Lemonade 11.8 Makes It Easy To Run DeepSeek V4 Flash On AMD Strix Halo - Phoronix
- Which AI brands have the most satisfied users? - YouGov
- Z.ai stock surges after launching GLM-5.3-Flash on Chinese chips - qz.com
- China open-source AI usage hits record high as DeepSeek leads - digitimes
- DeepSeek’s New Open Source AI System Modifies Its Own Code - Geeky Gadgets
- The Sequence Learning Loop - Issue #921: Learn About DeepSeek New Model, the Env Harness Paper and the Amazing Etched - TheSequence | Jesus Rodriguez
- Qwen & Zhipu AI Open-Source New Large Language Models Overnight – Both Priced Lower Than DeepSeek, Sparking Fierce Price War in Domestic Chinese LLM Market - 36 Kr
- EMXETF Launches China AI Tigers LLM ETF (NASDAQ: TGRZ) to Tap into China’s Leading AI Models - Yahoo Finance
- Open-Source AI Models Shift: Are Free Universal Uses No Longer Permitted for All Users? - 36 Kr
- Why open-weight AI is the new frontier in the US-China tech race - KrASIA
- Cantonese AI Model Pioneering Local Language AI Solutions - The Cryptonomist
- DeepSeek V4 Pro Safety Depends on the Agent Harness - quasa.io
- Zhong Shanshan Makes Rare AI Foray: 350 Million Yuan Indirect Bet on DeepSeek - finance.biggo.com
- MiniMax's Revenue Nearly Quadrupled While Its Losses Kept Growing Too - Startup Fortune
- Grok 4.6 on Microsoft Foundry - X.ai
- Elon Musk’s xAI Sounds Alarm Over Power Shutdown: Grok Could ‘Largely Cease to Function’ - Yahoo Finance
- Grok Bot is now included with more plans - X.ai
- Former sexual abuse victims say Grok used their images, videos to train deepfake capabilities - CyberScoop
- ChatGPT Voice vs Grok vs Gemini Live: $4.80/Hr Gap [2026] - tech-insider.org
- Federal suit filed in Western District of Arkansas: Grok created sexually explicit images of 17-year-old girl - Northwest Arkansas Democrat-Gazette
- NIU student used AI model Grok to create child sexual abuse material, authorities allege - Shaw Local
- Grok Can Now Build Custom Games: What You Need to Know - BASENOR - Tesla Accessories
- How to Use the Grok API: 12 Steps, 90 Min [2026] - tech-insider.org
- Claude Opus 5 vs Grok 4.6 vs Gemini 3.1 Pro: $19 Gap [2026] - tech-insider.org
- Grok Linear Triggers: What They Do and How to Set Them Up - BASENOR - Tesla Accessories
- Former xAI CFO Offers $100K Bounty for Info on Alleged Stalking That Began Hours After Abrupt xAI Departure - Glitchwire
- Tesla goes local in China, swapping Grok for ByteDance's Doubao LLM in smart cockpits - digitimes
- Tesla Summer Update Wires In Grok Voice Think Fast 2.0 - BASENOR - Tesla Accessories
- JPMorgan doubles down on SpaceX verdict on key update - thestreet.com
- New method systematically generates graph XAI benchmarks using Weisfeiler–Leman coloring - Bioengineer.org
- Rifles okay please: Yemen security check visuals spark internet buzz, X users include Grok | Watch | World News - Hindustan Times
- Codex Vs OpenCode Vs Pi Agent: Who Wins? Giuseppe Conte (Kdxb4bO2er) - Mshale
- AI向けの広告を9モデルで測ったら、採用したAIは一度も「広告だ」と言わなかった
- Qwen 3.8 Flash Next の N-GRAM/PLEを公式資料とDay zero実装から読み解く
- デザインガイド1枚で+12点、ローカルLLM生成の実測
- ローカルLLMをオーケストレーションしてサイトを作らせてみた
- Claude Code の Skill の description を最適化したら、動いたのは「発火率」じゃなかった
- 賢い方に道を作らせる:ローカルLLMを含んだシステムの詰まない設計
- AIに解雇された開発者が逆襲!オープンソースAI CEO「OpenExecutive」
- ドラフトトークンを並列生成する拡散モデル方式「DFlash」をGemmaで試した
- 改めてコンテキストエンジニアリングとは何か整理する
- Google Colabで最新LLMを試す #12 ― Nanbeige4.2-3Bを無料版T4で動かす:Tool Callingまで試す
- AIエージェントの暴走を止めるガード実装パターン集 — 実機ログから抽出した7つの型
- AIのプロンプトキャッシュ、当たらなくても+25%課金される — Azure GPT-5.6実測
- Claude Code が「言ってもいない自分の発言」を捏造した
- 自然言語ハーネスは素のレビューをどこまで変えるか:ChatGPT・Gemini・ClaudeでZIPあり/なしを比較した
- 【Unity】ねぇねぇみんな聞いて!! LLMをゲームAIに組み込んでみたよ!!!
- 社内スキルが190個になった日、エージェントに渡すべきは検索ではなく地図だと気づいた
- Claude API の自動キャッシュを使ったら、キャッシュなしより入力の料金が25%高くなりました
- Agent・Orchestration・Harness・Loop・Context・Evals、6つの関係を1つの具体例で整理する
- モデル切り替えは節約にならない——770セッションのログで測ったら会話の86%を読み直していた
- star 3.7万のAI Slop検出Skillは、日本語だと6パターンが空振りする
- Seq2Seq LSTMはランダムウォークを超えられるか | 第2回:AIで為替の未来予測は本当にできるのか?
- Attention機構とは?注目箇所を重み付けする仕組み
- fit_predictで体で覚えた、異常検知の定石
- Whisperの精度改善、5施策を実測したら3つ棄却になった話
- ニューラルネットは陰陽を学べるか。450エポックの間、目玉の正解率は0%だった
- BigQuery ML × Gemini でEC顧客の購買予測モデルを構築する
- AIって、「時間がもったいないとか」思うのかな?|医療AI・実践編 ⑧⏰
- M5 Ultra発表を機に、AI動画のピークメモリを見積もるPython
- Kimi K3を理解する②──アーキテクチャの全体像
- RAGは「一回検索」では足りない
- なぜAIは、専門家でなくても使えるようになったのか
- プロンプトを打つ時代から、ループを設計する時代へ
- 階層型言語モデル PHOTON を論文から実装し、RTX4090で再現できたこと・できなかったこと
- 因果推論 Day 14/全30回 反事実の計算、SCMで「もしあの時」を解く
- LLMのTrainingとInferenceを「検量線を引く・使う」で理解する
- イーロン・マスクとジャック・ヴァレ(AI対話)
- なぜ「LLMはAGIに到達するか?」という問いが無意味なのか
- AIエージェントを哲学から考える――八つの問いと八人の研究者
- Ornith-1.0-9B-MXFP4_Hybrid-Imatrix 総合ベンチマークレポート(全52問)
- シンガポールにおけるAIと著作権 / AI学習例外の安全弁の置き場 / robots.txtの評価が分かれる理由 / 発明者確定後に残る問い 雑感
- 1秒でタイムアウトするMCPクライアントに3秒かかるサーバーを繋いだら、2回目の質問に1回目の答えが返ってきた
- AIと金融 / 授権された取引と責任の分配 / OECD金融詐欺報告書 雑感
- AIリスク保険が吸収するもの / 流動性の時間差と保険法上の摩擦 雑感
- 電子の海でAIを叫ぶような話
- 日本のAI利用率は58.8%で、6.4%です。どちらも本当の数字なので、一次資料まで確かめました
- 推奨火力より少ない火力トークン量で効率よく高難度を解かせる ── SphereOS-Atlantis DOSが引き出す、財布(と地球)に優しいAIエンジニアリング
- 日誌(2026.08.24)
- 話題のグラフエンジニアリングについて、これまでの〇〇エンジニアリングを追ってみた
- ChatGPTで止まっている人へ。今読むべきClaude本おすすめベスト6
- AIも「寝かせる」と覚えがよくなる Google最新論文が描く"眠って夢を見るAI"の全貌【初心者向け徹底解説】
- AI↔キャラの切り替わりおもろい/会話ログ
- 🔊音声あり(日&英):【プロ直伝】AI動画広告の未来を変える!6つの評価ポイントを徹底解説【最新論文】
- 道具に纏わり付くAIという幻影 :思索スケッチ
- 【生成AIニュース+】『Muse Image』『Gemini 3.5 Transcribe』『Breeze TTS 2』『Ollama v0.33』『GLM-5.3-Flash』『H3 Max』『MiniMax-H3 Turbo』『H3-Optimizations』『MiniMax-H3-Acc-LoRAs』『MiniMax-H3 Experimental LoRAs』『MiniMax-H3 Director Cut Studio』他
- LLM#8 temperature 0.7 は、100回中99回なにもしなかった。効く場所が、198文字に1か所しかなかった
- LLM のプロンプトに「書くな」と書いても書く──自動記事パイプラインで踏んだ5つの穴
- MCPの新ロードマップ公開、今後はAIエージェント対応、HTTP通信への統一、アイデンティティ、よりよいデベロッパー体験などに注力
- AI活用率100%のQA組織をつくるまで
- LINEヤフーのAgent iを支えるAIエージェント基盤:「誰でも作れる」と「安全に動かせる」をどう両立したか
- 配属されて3か月。AI社員のBerry、現場で実際どんな仕事してるの?