AI News Digest 2026-09-12
特集
開発者コーナー
中堅コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧を表示
- DeepSeek releases V4.1-Flash [imp:75 dev:80]
- Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark [imp:45 dev:50]
- DeepSeek Taps CITIC Securities for Potential Shanghai IPO - Channel Insider [imp:15 dev:0]
- Show HN: Model pricing board for DeepSeek Harness: 7k models, cheapest route [imp:40 dev:75]
- A Human Audit of OpenAIs AI-Generated Mathematical Proofs [imp:55 dev:50]
- Former VP aide attacked AI rivals while holding $1M-plus stake in Elon Musk’s xAI - upnorthlive.com [imp:15 dev:5]
- Grok Bot Galaxy [imp:30 dev:60]
- SpaceXAI Says Alleged Grok Child Porn Maker Must Indemnify It - Law360 [imp:15 dev:0]
- ChatGPT最上位プラン(月額200ドル)が新規受付停止!「Astra」爆誕による影響と現在の裏側 (いいね相当スコア: 取得失敗) [imp:50 dev:15]
- OpenAI's new Agents API gives developers the infrastructure behind Codex and ChatGPT [imp:80 dev:80]
- OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time [imp:75 dev:80]
- v2.1.268 リリース - 価格統合とバグ修正 [imp:40 dev:70]
- Suno v6 [imp:35 dev:20]
- Quoting huggingface.co/security.txt [imp:5 dev:30]
- Native is now the future of mobile at Shopify [imp:40 dev:75]
- Soft-deprecating re.match() [imp:30 dev:75]
- Show HN: Toast, a beautiful by default in terminal IDE [imp:30 dev:75]
- Λ Snap – An inviting programming language for kids and adults for CS study [imp:15 dev:50]
- Litelm: LiteLLM Without the Bloat [imp:50 dev:75]
- Claude is only available to people over 18 years [imp:15 dev:10]
- The EPA Is Planning to Scrap Public Review Rules for Data Center Pollution [imp:50 dev:10]
- 118M Queries per Second on Neki [imp:60 dev:75]
- Zep AI (YC W24) Is Hiring a Head of Forward Deployed Engineering [imp:20 dev:75]
- Show HN: Godot and Rust based multiplexer (terminal panes and more) [imp:40 dev:75]
- Global Glacier Extinction Explorer [imp:10 dev:10]
- Show HN: Hacker News, Without AI [imp:25 dev:45]
- RTK reports token savings, but our cost benchmarks disagree [imp:60 dev:75]
- Measuring the sloppiness of code [imp:60 dev:75]
- Copying login keychains between Macs fails on Secure Enclave Macs with Tahoe [imp:40 dev:65]
- Show HN: Bodily Oddities [imp:5 dev:0]
- Room 641A [imp:5 dev:0]
- I've made an offer to buy Ryan Phelan and Stewart Brand's houseboat Mirene [imp:0 dev:0]
- Re-Engineering YouTube for the Living Room: Bringing "Chrobalt" to RDK [imp:40 dev:75]
- Cursor Projects - 大規模作業管理の新機能 [imp:75 dev:75]
- Rapidly scaling online storage to serve over 1 billion ChatGPT users [imp:65 dev:75]
- How a researcher uses Codex and ChatGPT to search for new antimicrobial molecules [imp:50 dev:50]
- Now everyone can put data to work [imp:60 dev:50]
- Introducing ChatGPT for Financial Services [imp:60 dev:50]
- Expanding AI access and cyber defense for federal, state, local, and tribal governments [imp:55 dev:10]
- 3 ways to prep for your next big race with Search [imp:5 dev:0]
- ToolGrad: Efficient tool-use dataset generation with textual "gradients" [imp:50 dev:75]
- Rebuilding AUTOMATIC1111 with Gradio Workflow [imp:50 dev:75]
- Marketing ops as code: Automating events from planning to follow-up on GitHub [imp:40 dev:75]
- GitHub Copilot app for Beginners: Using the diff, terminal, and browser [imp:40 dev:75]
- GitHub availability report: August 2026 [imp:25 dev:50]
- Introducing automatic remediation policies with Cloudflare CASB [imp:60 dev:75]
- 1.1.1.1 now supports post-quantum DNSSEC, all 2,420 bytes of it [imp:55 dev:75]
- Join Us at the Zephyr Project Meetup in Amsterdam [imp:10 dev:25]
- How Full-Stack NIM Optimizations Deliver 2.5x More Users on Nemotron 3 Ultra [imp:65 dev:75]
- From Wafer-Out to First Token: Codifying Supply Chain Expertise with Nemotron and Palantir Foundry [imp:65 dev:70]
- Powering AI is an architecture problem [imp:70 dev:65]
- Panic builds over bankrupt Spirit’s looming data sale to Google [imp:45 dev:10]
- An Anthropic researcher’s doomsday warning comes at a very interesting time [imp:55 dev:10]
- Nscale adds former OpenAI exec Fidji Simo to its board ahead of potential IPO [imp:25 dev:5]
- Jensen Huang explains why Nvidia will grow an astounding 70% next year [imp:35 dev:10]
- Mark Wahlberg is coming to TechCrunch Disrupt 2026, and he wants to talk about your work, not his [imp:5 dev:0]
- Meta’s AI agent Muse is now the No. 2 app in the US [imp:55 dev:25]
- India’s Pocket FM doubles revenue run rate to $500M as AI powers 93% of audio content [imp:55 dev:45]
- AI agents are flooding public services with new requests [imp:45 dev:10]
- Maven Robotics wants to steal your robot deployment deal [imp:45 dev:50]
- AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks [imp:35 dev:10]
- Meta says it’s changing AI suggestions after posing invasive personal questions [imp:40 dev:50]
- Slack can now vibe-code interactive charts and reports inside chats [imp:50 dev:50]
- Schools are catching on to Big Tech’s playbook [imp:35 dev:10]
- Universal Music is launching an AI music platform with ElevenLabs [imp:50 dev:20]
- Meta’s Muse AI works and creeps me out [imp:35 dev:25]
- Why the current tech backlash feels different [imp:35 dev:10]
- NVIDIA Personal AI Router Distributes AI Tasks Across Local Compute [imp:65 dev:75]
- How LinkedIn Trains AI Job Search 8x Faster with Multi-Teacher Distillation [imp:60 dev:75]
- Session Traces and Cost Controls Help Diagnose AI Agent Failures [imp:60 dev:75]
- OpenAI Releases GPT-6 Astra for Coding and Computer Use [imp:85 dev:75]
- Article: When Spec-Driven Development Pays Off [imp:60 dev:75]
- Ex-Deepmind VP Vinyals says AI self-improvement is coming but won't trigger an intelligence explosion [imp:45 dev:20]
- Deep learning pioneer Bengio argues the training process itself makes AI dangerous [imp:45 dev:15]
- OpenAI floats a shared AI slowdown, takes it to Congress [imp:70 dev:15]
- Class action lawsuit accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers [imp:35 dev:10]
- The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable [imp:45 dev:20]
- Anthropic's $1.5 billion book settlement descends into chaos as authors and publishers fight over who gets paid [imp:40 dev:10]
- OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows [imp:60 dev:75]
- Adaptive Entangled Game Modules in Artificial General Intelligence [imp:35 dev:30]
- Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks [imp:70 dev:75]
- Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions [imp:30 dev:30]
- An Autonomous GeoAI Agent for Arctic Eco-Navigation [imp:50 dev:70]
- The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents [imp:65 dev:75]
- Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery [imp:50 dev:50]
- Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration [imp:40 dev:50]
- XAI-Arena: Can LLMs Assess the Quality of XAI Explanations? [imp:50 dev:75]
- Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations [imp:60 dev:75]
- ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance [imp:60 dev:75]
- Multi-Agent Agentic Graph Learning via Structural Signatures [imp:55 dev:70]
- CityPlanner: A Sandbox Agent for Executable Urban Planning [imp:50 dev:65]
- A Function-Space Approach to the Statistical Mechanics of Learning Dynamics [imp:25 dev:20]
- From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins [imp:50 dev:65]
- Seven Sources of Physical AI Capability Formation [imp:55 dev:50]
- RobustSGPO: Search-Space Control for Agent Harness Evolution [imp:65 dev:75]
- Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery [imp:70 dev:75]
- RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems [imp:40 dev:50]
- PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations [imp:50 dev:75]
- Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents [imp:45 dev:50]
- Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation [imp:40 dev:70]
- Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning [imp:55 dev:80]
- Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety [imp:25 dev:55]
- LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents [imp:65 dev:85]
- Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks [imp:60 dev:80]
- Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward [imp:80 dev:85]
- UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model [imp:60 dev:85]
- The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents [imp:60 dev:75]
- Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields [imp:20 dev:50]
- AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents [imp:80 dev:85]
- Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format [imp:55 dev:75]
- Decision Transformer for UAV-Mounted RIS-Assisted Dynamic D2D Communications [imp:5 dev:15]
- Grounded Evaluation and Repair for NL-to-PDDL Problem Generation [imp:55 dev:80]
- Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models [imp:60 dev:80]
- Structural Process Supervision for Latent Chain-of-Thought Reasoning [imp:65 dev:85]
- Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability [imp:75 dev:85]
- OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization [imp:25 dev:55]
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States [imp:65 dev:80]
- RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases [imp:60 dev:80]
- Agent-Based ML-LLM Fusion with Self-Optimizing Prompts for Plateau Weather Alerts [imp:25 dev:50]
- Kernel-Managed Shared Memory for System-Wide Personalization [imp:75 dev:85]
- Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning [imp:60 dev:80]
- Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection [imp:55 dev:80]
- What Should an Agent Forget? Separating What Is Stored from What Is Used [imp:75 dev:85]
- TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards [imp:75 dev:85]
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning [imp:60 dev:80]
- Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System [imp:80 dev:55]
- Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs [imp:75 dev:85]
- ConvMem: Convolutional Memory for Long-Context Reasoning [imp:75 dev:85]
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition [imp:75 dev:85]
- Quantifying Logical Consistency in Transformers via Query-Key Alignment [imp:60 dev:80]
- From Plausible to Actionable: A Position on LLM Self-Explanations [imp:60 dev:75]
- Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling [imp:20 dev:55]
- Trust Me, I'm Your Developer: Self-Issued Authentication in Large Language Models [imp:65 dev:75]
- AgenticGen: Reward-Guided Agentic Video Generation for Advertising [imp:25 dev:50]
- Reliability-Aware Hybrid-K Ensemble Selection for Cervical Cytology Classification: Integrating Discrimination, Calibration, and Selective Prediction [imp:20 dev:55]
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents [imp:80 dev:80]
- Geometry Conditioning in an Embodied SLM: Training Controls and Robustness Diagnostics in a 0.8B Hybrid Model [imp:60 dev:80]
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents [imp:75 dev:80]
- Compute-Bounded Security Assurance - Coverage, Verification, and Response under Resource Constraints [imp:60 dev:75]
- Scaling Post-Training Ternarisation to Qwen3-8B Capability Retention, Reproduction, Lossless Packing, and Packed Execution [imp:75 dev:85]
- Distribution-Consistent Inference for Dynamic Sparse Mixture-of-Experts [imp:75 dev:85]
- Talking to Itself While Coding: What Makes Comments Help Code Generation? [imp:60 dev:80]
- In RAG We Trust? Measuring Robustness of Retrieval-Augmented Generation Under Document Poisoning [imp:75 dev:80]
- Critical initialization destabilizes higher input derivatives in wide scalar-input networks [imp:5 dev:50]
- What Fixed-Rollout pass@k Evaluations Can Identify [imp:55 dev:75]
- No Free Checker: A Survey of Verifiers for Robot Policies [imp:55 dev:80]
- DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity [imp:60 dev:80]
- Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models [imp:60 dev:75]
- Support Discovery With Iteratively Reweighted Least Squares for Fixed-Charge Network Flow [imp:20 dev:75]
- Improving 5G AI-RAN MCS Selection by Predicting Retransmissions [imp:15 dev:50]
- Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management [imp:65 dev:80]
- Auditable Emergency Triage for Maternal and Newborn Care in India [imp:60 dev:50]
- VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models [imp:75 dev:80]
- An Experimental Evaluation of Multimodal Prompt Injection Attacks on Agentic AI Frameworks [imp:75 dev:80]
- Reliable Near-Field Multi-User Positioning Informed by Two-Stage MUSIC [imp:15 dev:50]
- Edu-QuRating: Multi-Dimensional Educational Data Curation with Distilled Pairwise Judgements [imp:55 dev:75]
- SCCM : Stream Cruise Control Method for Automated Drift Detection and Adaptation [imp:60 dev:80]
- Efficient Leakage-Free Neural Architecture Search under Leave-One-Subject-Out Evaluation [imp:60 dev:80]
- Distributed Physical Layer Authentication and Collaborative RSMA in Non-Terrestrial Networks via Graph Reinforcement Learning [imp:20 dev:55]
- From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls [imp:60 dev:80]
- Adaptive Distributed Physical-Layer Authentication and Attack Detection in 6G Non-Terrestrial Networks via Causal Meta-Learning [imp:20 dev:55]
- A Statistical Approach to Estimating Sample Size of Machine Learning Models [imp:60 dev:80]
- Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning [imp:75 dev:80]
- High-probability guarantees for linear accessibility in feature superposition [imp:5 dev:50]
- The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption [imp:65 dev:85]
- Learning with Synthetic Data via SGD in High-Dimensional Linear Regression [imp:20 dev:75]
- Myocardial Strain Drift Correction in Deep Learning Based Ultrasound Tracking [imp:20 dev:55]
- Modality-Decoupled Federated Learning for Privacy-Preserving Embodied Intelligence in 6G [imp:65 dev:80]
- Teacher Geometry Shapes Learnability in Teacher-Student Networks [imp:20 dev:75]
- Compact Visuotactile World Models for Lifting: Prediction, Reward Alignment, and Force Constraints [imp:60 dev:80]
- Watermarks Without Verification: AI Text Watermarking After the EU AI Act [imp:80 dev:50]
- RouteBridge: Reliability-Routed Bidirectional Distillation Between Neural Radiance Fields and 3D Gaussian Splatting [imp:60 dev:80]
- Hyperbolic Geometry for Open-World Object Detection in Remote Sensing Imagery [imp:60 dev:80]
- Cascading Gradient Inversion via LT-Code Inspired Peeling in Federated Learning [imp:75 dev:80]
- Introducing Consort: A Spec-First Agent Framework for Enforced, Test-Driven Development on Live Database Branches [imp:75 dev:90]
- Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA [imp:55 dev:55]
- Looped GPT-BERT: Trading Parameters for Computation in Small Language Modeling [imp:55 dev:80]
- CT-SAFR: Safe and Interpretable Chain-of-Thought Reasoning for Autonomous Robots: A Multi-Layered Verification Framework for Trustworthy AI-Driven Robotic Decision Making [imp:60 dev:80]
- When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination [imp:75 dev:80]
- Kernel-Complexity Edge Sanitization for Training-Free Defense against Structural Graph Attacks [imp:60 dev:80]
- Distilling Image Prototypes for Guided Test-Time Adaptation [imp:60 dev:80]
- HiRAD: A Flexible Large-Scale AGV Routing System [imp:60 dev:80]
- Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both? [imp:75 dev:85]
- BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL [imp:75 dev:85]
- Pairit: A Platform for Live Experiments on Human-AI Collaboration [imp:65 dev:85]
- LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios [imp:60 dev:80]
- How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE [imp:75 dev:85]
- CS-Guard: Benchmarking LLM Guardrails for Code Generation Security [imp:75 dev:85]
- uFlowCSP: Crystal Structure Prediction using Mean flow generative models [imp:25 dev:80]
- Subgroup Membership Inference Audits of Differentially Private Synthetic Text [imp:75 dev:80]
- Can AI Agents Detect and Repair Artifact Drift in Network Experiments? [imp:75 dev:85]
- With a Thermomix You Lose the Ability to Cook: A Kitchen Machine Analogy for Applications of Generative AI in Education [imp:5 dev:20]
- Forward-Free LLM Depth Pruning via Weight Redundancy [imp:75 dev:85]
- Albedo Estimation via Latent Bridge Matching [imp:20 dev:80]
- Strangers to Themselves: What Language Models Say About Themselves Is Generic [imp:60 dev:80]
- FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models [imp:75 dev:85]
- Improving Cross-Lingual Token Representations by Adding a Pinch of SALT [imp:60 dev:80]
- Fidelity-Aware Scheduling of Quantum Circuits on Multi-QPU Systems [imp:20 dev:55]
- What Makes Adversarial Examples Transfer Across Deepfake Detectors? [imp:75 dev:80]
- MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes [imp:60 dev:80]
- Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation [imp:50 dev:60]
- Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training [imp:55 dev:75]
- NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments [imp:40 dev:65]
- A statistical approach to bias in zero-shot learning: the lens of handwriting recognition [imp:30 dev:55]
- Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance [imp:70 dev:70]
- A Trust-Network-Based Federated Learning Framework for Multi-Center Aging Clock Prediction [imp:35 dev:50]
- SA-Profile: Automated Sulcus Angle Profiling from Super-Resolution MRI [imp:15 dev:30]
- Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design [imp:50 dev:75]
- Active Adaptation, Not Static Defense: Temporal Dynamics of Preventative Steering in Adversarial Fine-Tuning [imp:60 dev:75]
- Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries? [imp:55 dev:80]
- Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search [imp:40 dev:60]
- LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation [imp:55 dev:75]
- A-JIT: Agentic Just-In-Time Software Construction [imp:65 dev:85]
- DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs [imp:45 dev:65]
- GANDR: Claim Auditing for Verifiable Legal Answer Generation [imp:50 dev:70]
- Learning Intrusion Response Strategies for OT Systems [imp:50 dev:65]
- RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding [imp:55 dev:80]
- One Loop, Two Gains: Can Active Learning win the Lottery for Free? [imp:45 dev:70]
- Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs [imp:55 dev:75]
- OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis [imp:45 dev:60]
- PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving [imp:50 dev:75]
- MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production [imp:30 dev:50]
- Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization [imp:50 dev:70]
- Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support [imp:20 dev:40]
- Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs [imp:60 dev:80]
- Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization [imp:50 dev:75]
- IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier [imp:55 dev:85]
- Show-Harness: Just a VLM Agent Can Play Robots [imp:55 dev:80]
- Reinforcement Learning with Temporal-Logic-Based Causal Diagrams [imp:45 dev:70]
- Reinforcement learning for Quantum Tiq-Taq-Toe [imp:15 dev:50]
- ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork [imp:45 dev:75]
- RelayS2S: A Dual-Path Speculative Generation for Real-Time Dialogue [imp:50 dev:75]
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction [imp:50 dev:70]
- Zero-shot World Models Are Developmentally Efficient Learners [imp:45 dev:65]
- Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies [imp:35 dev:75]
- CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification [imp:40 dev:65]
- Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations [imp:60 dev:80]
- Cultural Binding Heads in Language Models [imp:50 dev:75]
- KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning [imp:50 dev:80]
- Self-Evolving Scientific Agent Designs Physically Reasoned White-Box Fluid Control [imp:55 dev:85]
- EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading [imp:50 dev:80]
- KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms? [imp:60 dev:90]
- ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion [imp:45 dev:70]
- Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning [imp:50 dev:75]
- LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting [imp:45 dev:80]
- Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation [imp:55 dev:85]
- Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents [imp:50 dev:80]
- FrontierChallenge: Evaluating Scientific Workflow Completion [imp:60 dev:85]
- A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation [imp:60 dev:85]
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation [imp:70 dev:90]
- Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses [imp:70 dev:90]
- From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale [imp:65 dev:90]
- DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning -- Extended Version [imp:35 dev:60]
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis [imp:65 dev:90]
- RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts [imp:30 dev:50]
- EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering [imp:50 dev:80]
- Equity Promotion in Online Resource Allocation [imp:30 dev:50]
- Incentives to Offer Algorithmic Recourse [imp:35 dev:55]
- A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decision Model [imp:70 dev:90]
- BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language Models [imp:55 dev:75]
- Influence-Oriented Personalized Federated Learning [imp:45 dev:75]
- Efficient Diversity-based Experience Replay for Deep Reinforcement Learning [imp:45 dev:75]
- Query Brand Entity Linking in E-Commerce Search [imp:35 dev:55]
- Safe Learning Under Irreversible Dynamics via Asking for Help [imp:50 dev:75]
- Predicting Estimated Times of Restoration for Electrical Outages Using Longitudinal Tabular Transformers [imp:30 dev:50]
- Synergistic Vision-Language Reinforcement Enables Scalable On-Demand Analysis across Diverse Clinical Tasks [imp:50 dev:70]
- SloMoDeblur: A Large-Scale Smartphone Image Deblurring Dataset [imp:25 dev:40]
- Instance-Aware Algorithm Selection for Maximum Clique via a Dual-Channel Graph Neural Architecture [imp:40 dev:75]
- RAU: Reference-based Anatomical Understanding with Vision Language Models [imp:45 dev:65]
- MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation [imp:50 dev:80]
- Generative AI for Analysts [imp:40 dev:50]
- Meta-RL with Bayesian Linear Task Models [imp:45 dev:80]
- From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges [imp:55 dev:85]
- Elsewise: Authoring Open-ended Interactive Narrative with Possibility Space Visualization [imp:35 dev:55]
- Toward Learning POMDPs Beyond Full-Rank Actions and State Observability [imp:45 dev:80]
- Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist [imp:45 dev:70]
- Revisiting the Shape Convention of Transformer Language Models [imp:50 dev:85]
- False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK [imp:35 dev:60]
- City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification [imp:45 dev:80]
- MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration [imp:70 dev:95]
- Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework [imp:55 dev:80]
- Spec-Harness: Measuring and Improving Behavioral Adequacy of LLM-Synthesized Formal Specifications [imp:60 dev:90]
- Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning [imp:50 dev:85]
- Where is the Mind? Persona Vectors and LLM Individuation [imp:40 dev:70]
- The Biggest Risk of Embodied AI is Governance Lag [imp:55 dev:70]
- Dont Just Teach, Explain! A Gamified 20Q Recommender for Cybersecurity Education [imp:25 dev:50]
- "What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use [imp:35 dev:60]
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents [imp:65 dev:90]
- Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs [imp:55 dev:85]
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents [imp:65 dev:90]
- Tracing Computation Density in LLMs [imp:50 dev:85]
- BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language [imp:25 dev:50]
- Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning [imp:45 dev:80]
- FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (Sep 3rd version) [imp:55 dev:85]
- FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning [imp:45 dev:80]
- Expert-Level Crisis Detection in Mental Health Conversations [imp:40 dev:65]
- PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement [imp:35 dev:60]
- FP8 is All You Need (Part 2): Full-FP64 3-D FFT on FP8-Generation Tensor CoresThe Integer-Epilogue Wall and the Minimal Hardware That Would Remove It [imp:50 dev:85]
- Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning [imp:30 dev:75]
- Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software [imp:70 dev:90]
- PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal T2 mapping [imp:25 dev:25]
- DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation [imp:55 dev:65]
- Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models [imp:35 dev:45]
- Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees [imp:30 dev:50]
- Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability [imp:50 dev:60]
- Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning [imp:65 dev:80]
- tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots [imp:30 dev:45]
- 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection [imp:50 dev:65]
- AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP [imp:40 dev:50]
- LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation [imp:40 dev:55]
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach [imp:20 dev:45]
- Phase-Aware Spatial-Frequency Fusion for Few-Shot Fine-Grained Image Classification [imp:20 dev:40]
- Influence of Extruded Filament Shape on Buildability in 3D Concrete Printing: A Geometry-Informed Deep Learning-FEM Approach [imp:15 dev:25]
- VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models [imp:50 dev:65]
- PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation [imp:40 dev:60]
- Programmable Cellular Automata [imp:15 dev:40]
- When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization [imp:15 dev:35]
- Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs [imp:55 dev:70]
- Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web [imp:25 dev:45]
- SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities [imp:20 dev:35]
- Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling [imp:40 dev:55]
- Omni Interaction Agent Technical Report [imp:0 dev:0]
- M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction [imp:20 dev:40]
- Halo: Improving forecast accuracy through heteroscedastic estimation [imp:30 dev:55]
- Zero-shot rib design: merging training-free generative prior with topology optimization [imp:35 dev:55]
- Byzantine-Robust Federated Fire Detection with a Rotating Coordinator [imp:35 dev:60]
- Artificial Intelligence Algorithms for the Detection of Pathologies Related to Lung Cancer through Image Analysis using Convolutional Neural Networks and Data Augmentation: a systematic mapping of the literature [imp:25 dev:35]
- GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models [imp:50 dev:70]
- Conformal Calibration Transfer [imp:35 dev:55]
- The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes [imp:25 dev:50]
- Adaptive Margin Ordinal Loss: Penalizing Center-Class Hedging in Ordinal Classification [imp:20 dev:45]
- A Bellman Optimality Equation for Plasticity [imp:30 dev:50]
- Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables [imp:35 dev:55]
- From Connectivity to Rewards: Dense Reward Learning with Directed State Graphs [imp:30 dev:50]
- DR-LabStack: Design and Implementation of a Clinician-Facing Web System for Diabetic Retinopathy Prediction [imp:30 dev:60]
- RiVaT-Fuse: Reliability-Calibrated Variational Tensor Fusion for Multimodal Prediction under Modality Uncertainty [imp:25 dev:50]
- Processing and classifying bird songs using wavelet techniques and supervised learning [imp:15 dev:35]
- Flow Duality and Source Geometry for Categorical Generation [imp:30 dev:50]
- Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations [imp:30 dev:55]
- Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry [imp:25 dev:45]
- Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble [imp:45 dev:60]
- Relatively Smart II: Tractable or Semi-Supervised Instance-Optimal Learning [imp:25 dev:45]
- AUC Maximization from Biased Positive-unlabeled Data with Confidence [imp:25 dev:50]
- Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation [imp:40 dev:60]
- When More Is Not Better: Component Anti-Synergy in a P300 Speller [imp:20 dev:40]
- Phases in a class of associative memories via hidden neurons [imp:25 dev:45]
- EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale [imp:45 dev:65]
- Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret [imp:25 dev:45]
- Importance Weighting for Unlabeled-unlabeled Learning under Distribution Shift [imp:25 dev:50]
- Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control [imp:35 dev:55]
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks [imp:75 dev:85]
- EMMI: Edge Multi-Modal Intelligence for Communication-Efficient MLLM Inference via Fused Representation Compression [imp:50 dev:70]
- The information geometry of large language models is shared, learned, and controllable [imp:50 dev:65]
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization [imp:50 dev:65]
- HERALD: High-Fidelity Exemplar Retrieval with Adaptive Landmark Distillation for Heterophily-Aware Graph Condensation [imp:30 dev:55]
- How Wrong Can a Good Predictor Be? Diverging Updates with Vanishing Predictive KL [imp:25 dev:45]
- Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving [imp:60 dev:75]
- Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models [imp:35 dev:55]
- LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry [imp:60 dev:75]
- When does a spectral prior help graph learning? Connectivity-loss estimation under road-network disruptions [imp:30 dev:55]
- Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications [imp:25 dev:45]
- Hierarchical Clustering Can Jointly Satisfy Richness, Consistency, and Scale Invariance [imp:35 dev:55]
- Convex Optimization with Nested Evolving Feasible Sets (CONES) under Time-Varying Loss Functions [imp:25 dev:50]
- REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving [imp:60 dev:75]
- Legible Failures: Detecting and Repairing In-Context Binding Errors [imp:45 dev:65]
- Polyhedral Geometry of Time-to-First-Spike Neural Networks [imp:30 dev:50]
- Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer [imp:30 dev:55]
- MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions [imp:45 dev:65]
- A Dynamic Fusion Large Language Model for Traffic Flow Prediction [imp:30 dev:55]
- Estimating Inconsistency Response Surfaces under Uncertainty in Cyber-Physical System Development [imp:25 dev:45]
- Reification as a Transferable Vocabulary: Zero-Shot Link Prediction with Vanilla GNNs [imp:40 dev:60]
- Local Robustness Quantification for Naive Bayes Classifiers and Generative Forests: a General Approach [imp:30 dev:55]
- Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets [imp:25 dev:45]
- Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study [imp:35 dev:55]
- DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis [imp:40 dev:60]
- Generalized Score Matching for Parameter Estimation on Convex Domains [imp:25 dev:50]
- Particle GFlowNets: Rethinking Generative Marginalization Models [imp:30 dev:55]
- A Dataset and Model for Imputing Water Surface Elevation on a Large and Extremely Sparse Spatiotemporal Graph [imp:20 dev:40]
- LoaDiff: Conditional Generation of Electricity Consumption Time Series for Energy Analytics [imp:35 dev:55]
- RDDMPI: Residual Denoising Diffusion Model for Probabilistic Multivariate Time Series Imputation [imp:40 dev:60]
- Musec: MomentUm SpEctral Clipping for Stable Muon-type Training [imp:60 dev:75]
- Learnware and AI Model Management System [imp:55 dev:75]
- Why Does Post-Training Quantization Work? [imp:60 dev:75]
- Predicting Privacy Leakage from Weight Spectral Density [imp:50 dev:65]
- Dynamic language model representations for multi-objective reaction optimisation [imp:30 dev:55]
- Thinking with Looped Flows [imp:40 dev:60]
- Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport [imp:40 dev:60]
- AdamX: Cosine similarity meets gradient descent [imp:45 dev:70]
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement [imp:50 dev:65]
- CoRA-NAS: Coarse Ranking and Anchor-Residual Refinement for Neural Architecture Search [imp:35 dev:60]
- CausalArena: Benchmarking Causal Discovery in the Foundation Model Era [imp:55 dev:70]
- TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription [imp:25 dev:50]
- From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good [imp:35 dev:40]
- Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data [imp:55 dev:70]
- General Quantification of Covariate and Concept Shifts [imp:50 dev:65]
- MUC-FL: Block-Wise Marginal Utility Contribution for Communication-Efficient Federated Learning [imp:45 dev:65]
- Optimizing AI Inference Across the Deployment Stack [imp:65 dev:80]
- EVTradeMatch: A Mobility-Aware Multi-Objective Matching Framework for EV--EV Energy Trading [imp:30 dev:55]
- Supply Chain Analytics: A Data-Driven Approach [imp:40 dev:55]
- A Station-Based Evaluation of Machine Learning-based Weather Forecasting Models in Northern Norway [imp:25 dev:50]
- An Empirical Measurement of Jailbreaking Evaluators [imp:40 dev:60]
- Adaptive Diffusion Freezing: Privacy-preserving Diffusion Models Against Membership Inference Attacks [imp:45 dev:70]
- On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study [imp:50 dev:80]
- Black-Box Membership Inference via Word-Level Probability Estimation [imp:55 dev:70]
- PEARL: A Task-Aware Framework for Evaluating Differentially Private Synthetic Educational Data [imp:35 dev:60]
- Understanding In-Context Multimodal Jailbreaks via Posterior Reweighting [imp:50 dev:70]
- SoK: Privacy Attacks on Machine Learning via Explainable AI [imp:60 dev:75]
- From Cycle Space to Cycle Manifold: Limits and Achievability of Blind False Data Injection Attacks [imp:20 dev:20]
- Numbat: Building and Verifying a Self-Contained Machine-Learning Stack [imp:65 dev:85]
- Sequence-Informed Geometric Evaluation of RNA 3D Structures [imp:30 dev:40]
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive Reasoning [imp:50 dev:65]
- Understanding LoRA Rank Trade-offs in Diffusion Model Fine-Tuning [imp:50 dev:75]
- Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking [imp:50 dev:70]
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining [imp:55 dev:70]
- A Quantum-Inspired Dequantization Method for Diagonally Weighted Matrix Functions: Application to Learning with Optimized Random Features [imp:30 dev:60]
- SynCo: Synthetic Community-Aware Attributed Graph Generator for Graph Neural Network Benchmarking [imp:40 dev:65]
- CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding [imp:35 dev:65]
- Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems [imp:45 dev:70]
- Meta-Learning for Data-Efficient Plant Growth Estimation via Vision Transformers and Fuzzy Clustering [imp:35 dev:65]
- When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents [imp:60 dev:85]
- Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu [imp:40 dev:50]
- Weighted Empirical Risk Minimization for Machine Learning under Long-Range Dependence: Exact Pathwise Rates and Learning-Error Geometry [imp:35 dev:60]
- Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving [imp:65 dev:85]
- How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark [imp:15 dev:15]
- Studying Without a Syllabus: Task-Agnostic Environment Preprocessing [imp:55 dev:80]
- Scale-Aware 3D Deep Learning for Robust Brain Metastasis Detection in Multimodal MRI [imp:50 dev:70]
- Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models [imp:55 dev:75]
- scDEFT: A deep learning framework for drug-effect prediction and counterfactual reasoning [imp:50 dev:70]
- Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions [imp:50 dev:70]
- Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data [imp:35 dev:55]
- DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents [imp:65 dev:80]
- Symmetry-aware super-resolution of crystal orientation maps via invariant latent-space learning [imp:35 dev:65]
- ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations [imp:55 dev:75]
- Structurally Speaking: Motif-Oriented Graph Captioning through Bidirectional Graph-Text Translation [imp:40 dev:65]
- Empirical Evaluation of Membership Inference Attacks on NLP Text Classifiers: A Baseline Study on SST-2 [imp:50 dev:70]
- The Platonic brain bridge hypothesis: human brain networks as an architectural prior for omni models [imp:50 dev:70]
- Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction [imp:35 dev:65]
- Testing Between the Test Cases: Proving End-to-End Steering in Conditions You Never Drove [imp:60 dev:80]
- Empirical Evaluation of Data Poisoning Attacks in Supervised Learning [imp:50 dev:70]
- A variational physics-informed graph neural network for heterogeneous solid mechanics [imp:40 dev:70]
- New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models [imp:45 dev:70]
- Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss [imp:50 dev:75]
- Meta-Learning for Classifier Selection in Image Datasets: A Feature-Driven Framework for Accuracy Prediction [imp:45 dev:70]
- When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text [imp:50 dev:75]
- Coherent Floquet quantum reservoirs for molecular property prediction [imp:35 dev:60]
- TailProp: content-adaptive light- and heavy-tailed propagation for vision [imp:45 dev:70]
- The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems [imp:60 dev:75]
- A Fragility Spectrum for Recursive Language-Model Training [imp:60 dev:75]
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series Forecasting [imp:40 dev:65]
- Diversity of EML-type operators [imp:10 dev:15]
- Generative Replay Mitigates Sample Starvation in Quantum Architecture Search [imp:35 dev:60]
- Rethinking Radiomap Blind Prediction with Limited Environment and Configuration Representations [imp:30 dev:60]
- Improving Faint Object Detection for Space Situational Awareness with Variational Autoencoders [imp:45 dev:70]
- Predicting Train Delays in Finland Using Machine Learning and Weather Data [imp:35 dev:65]
- Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 1 [imp:45 dev:70]
- A Hilbert-Valued Functional Decomposition Framework for Explaining Time-Dependent Outputs [imp:40 dev:70]
- A Two-Mirror Faceted Projection System for EUV Lithography [imp:15 dev:15]
- Your Model Already Knows Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models [imp:55 dev:80]
- E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets [imp:35 dev:60]
- Deep operator learning for efficient sampling from invariant measures of stochastic differential equations [imp:40 dev:70]
- VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents [imp:65 dev:85]
- Improving the Sensitivity of Gravitational Wave Detection with Weighted Conformal Prediction [imp:50 dev:75]
- Hologram Representation via Quadratic Phase Gaussian Splatting [imp:45 dev:70]
- Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints [imp:55 dev:80]
- Structural priors for data-efficient language learning [imp:50 dev:75]
- Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution [imp:45 dev:70]
- Risk-Averse Decision Making with Multi-Level Reliability Guarantees [imp:40 dev:70]
- Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN) [imp:60 dev:80]
- A distribution-free certification framework for trustworthy crash-severity prediction [imp:50 dev:75]
- Identifiability of Nonnegative Tensor Decompositions via Positive Scattering [imp:35 dev:65]
- Distributed Optimization of Modular Production Systems using Model-based Reinforcement Learning with Inverse Models [imp:50 dev:75]
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation [imp:70 dev:80]
- ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding [imp:60 dev:80]
- Multimodal Taxonomic Conditioning for Generative Plankton Imagery [imp:45 dev:70]
- Geospatial Foundation Models Capture Health-Relevant Dimensions of Place Beyond Conventional Social Risk Indices [imp:55 dev:75]
- Negative Self-Distillation: Learning to Reason by Avoiding Flaws [imp:60 dev:80]
- Generalization Analysis of Distributed Kernel-based Robust Gradient Descent Algorithms [imp:40 dev:70]
- Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion [imp:50 dev:75]
- Learning structural balance of graphs from quantum spectral features [imp:35 dev:60]
- ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI [imp:60 dev:80]
- LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation [imp:65 dev:85]
- Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs [imp:65 dev:85]
- Sparsity Regularized and Robust Mean Variance Portfolio Selection Under Ellipsoidal Uncertainty [imp:35 dev:65]
- SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control [imp:60 dev:80]
- A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients [imp:50 dev:80]
- Differentially Private EEG Feature Anonymization: A Privacy-Utility Case Study in Clinical Neurophysiology [imp:50 dev:75]
- Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling [imp:55 dev:80]
- Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead [imp:50 dev:75]
- Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models [imp:50 dev:75]
- Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting [imp:55 dev:80]
- Domain-Specific Hallucination Detection in Large Language Models [imp:60 dev:80]
- 3D Point Splatting for mmWave Radar Novel View Synthesis [imp:45 dev:75]
- Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact [imp:55 dev:75]
- Optimizing Three Critical Factors for Practical and Effective OOD Detection Fine-Tuning [imp:50 dev:80]
- Label Differential Privacy via Aggregation [imp:50 dev:75]
- DNA: Differentially private Neural Augmentation for contact tracing [imp:45 dev:75]
- ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks [imp:55 dev:85]
- Mapping Seven Decades of Philosophy in Colombia: Dynamic Topic Modelling of Ideas y Valores [imp:25 dev:40]
- SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations [imp:50 dev:80]
- Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis [imp:50 dev:75]
- CertDW: Towards Certified Dataset Ownership Verification via Conformal Calibration [imp:25 dev:35]
- Learning Intrinsic Water-Quality Dynamics with Rainfall for Data-Driven Forecasting [imp:10 dev:15]
- Evidence for Limited Metacognition in LLMs [imp:45 dev:55]
- BiHDTrans: binary hyperdimensional transformer for efficient multivariate time series classification [imp:35 dev:50]
- On the Societal Impact of Machine Learning [imp:35 dev:45]
- Autonomous-Flow-Based Generation [imp:30 dev:40]
- When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification [imp:35 dev:55]
- UBCL: A Reinforcement Learning Framework for Controllable and Diverse Player Behaviors [imp:25 dev:45]
- Output Embedding Centering for Stable LLM Pretraining [imp:50 dev:65]
- Semidefinite Programming for Quantum Channel Learning [imp:20 dev:25]
- Smoothing the Score Function to Enhance Generalization in Diffusion Models [imp:45 dev:55]
- Prediction--Loss Alignment for Sampler--Robust Flow Matching Training [imp:40 dev:60]
- Evaluating Memory Structure in LLM Agents [imp:50 dev:70]
- Partial GFlowNet: Accelerating Convergence in Large State Spaces via Strategic Partitioning [imp:35 dev:45]
- Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification [imp:60 dev:70]
- Longitudinal Risk Prediction in Mammography with Privileged History Distillation [imp:30 dev:35]
- HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention [imp:55 dev:75]
- FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving [imp:60 dev:75]
- Generalization Guarantees on Data-Driven Tuning of Gradient Descent with Langevin Updates [imp:35 dev:50]
- Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations [imp:30 dev:40]
- DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum [imp:45 dev:60]
- Assessing Predictive Models for Fairness Based on Activity-Space Patterns [imp:35 dev:50]
- Benchmarking non-conformity score functions in conformal prediction [imp:30 dev:45]
- Using Seismic Statistical Features and VQ-VAE to Improve Spatiotemporal Seismicity Predictability [imp:25 dev:30]
- Attention by Synchronization in Coupled Oscillator Networks [imp:35 dev:45]
- SafeImpute: Reliable Clinical Data Imputation via Conformal Selection [imp:35 dev:45]
- RDQ: Residual Distribution Quantization for Large Language Models [imp:55 dev:75]
- Progressive Agent Skill Generation via Reinforcement Learning [imp:65 dev:80]
- Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction [imp:20 dev:25]
- Scaling Automatic Research Agents via World Models [imp:65 dev:75]
- Transfer Learning of Keystroke Dynamics for Cross-Device User Authentication [imp:30 dev:40]
- Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment [imp:55 dev:70]
- Toward a First-Principles Update Geometry for the Language-Model Head [imp:50 dev:70]
- Revenge of Monosemanticity: Neuron Specialization as a New Form of Feature Learning in MLPs [imp:45 dev:55]
- CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery [imp:40 dev:60]
- How Proper Scoring Rules Shape LLM Forecasting [imp:45 dev:60]
- REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent [imp:55 dev:75]
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory [imp:60 dev:75]
- Positional task conditioning for scalable defect detection across product families in large product catalogs [imp:50 dev:65]
- Time-Varying Graph Learning with Constraints on Graph Temporal Variation [imp:30 dev:45]
- Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients [imp:30 dev:45]
- No Screening is More Efficient with Multiple Objects [imp:15 dev:10]
- Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters [imp:40 dev:55]
- The observational partial order of causal structures with latent variables [imp:30 dev:40]
- Quantum State Preparation with the QNN-based SRBB Algorithm [imp:20 dev:25]
- Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks [imp:25 dev:40]
- Divergence-Based Similarity Function for Multi-View Contrastive Learning [imp:35 dev:55]
- Test time training enhances in-context learning of nonlinear functions [imp:45 dev:65]
- Configuration-Dependent Lower Bounds for Approximation by Shallow ReLU$^k$ Networks on the Sphere [imp:20 dev:30]
- Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge [imp:45 dev:60]
- PitchFlower: A flow-based neural audio codec with pitch controllability [imp:40 dev:55]
- Addressing A Posteriori Performance Degradation in Neural Network Subgrid Stress Models [imp:35 dev:45]
- Statistical analysis of Inverse Entropy-regularized Reinforcement Learning [imp:30 dev:45]
- Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts [imp:45 dev:60]
- Building Supervision into Hebbian Plasticity through Spike Agreement [imp:30 dev:40]
- Bayesian quantum sensing using graybox machine learning [imp:25 dev:35]
- Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids [imp:30 dev:40]
- Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret without a Variance Bound [imp:15 dev:10]
- mmFHE: mmWave Sensing with End-to-End Fully Homomorphic Encryption [imp:45 dev:60]
- Perturbation: A simple and efficient adversarial tracer for representation learning in language models [imp:45 dev:60]
- Active noise cancellation on open-ear smart glasses [imp:35 dev:45]
- Wiggle and Go! System Identification for Zero-Shot Dynamic Rope Manipulation [imp:40 dev:55]
- Discriminative Span as a Predictor of Synthetic Data Utility via Classifier Reconstruction [imp:45 dev:60]
- Continuous Diffusion Scales Competitively with Discrete Diffusion for Language [imp:50 dev:70]
- Goal-Oriented Lower-Tail Calibration of Gaussian Processes for Bayesian Optimization [imp:40 dev:60]
- CLSP-REQA: A Real-Time Quality-Aware Closed-Loop Seizure Prediction Framework with Mamba-BiLSTM and Confidence-Gated Intervention [imp:50 dev:60]
- Activation-Based Active Learning for In-Context Learning: Challenges and Insights [imp:50 dev:70]
- AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models [imp:60 dev:75]
- Statistically Valid Post-Training Hyperparameter Selection: From Tuning to Guarantees [imp:45 dev:65]
- Physics-constrained neural networks for surrogate modeling of lossless periodic structures [imp:40 dev:55]
- ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification [imp:45 dev:60]
- The Zero Pattern of a Design Matrix Drives Multiple Descent in Over-parameterized Regression [imp:30 dev:45]
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups [imp:50 dev:65]
- SparseDitto: An Agentic Sparse Compilation Framework through Architecture-Aware Synthesis on GPUs [imp:55 dev:75]
- VALG: An Agentic System for ML Theory Research [imp:60 dev:75]
- What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation [imp:40 dev:55]
- Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents [imp:65 dev:80]
- SAC-Copula: Quality-Preserving Watermarking for Diffusion Language Models via Smooth Correlated Gumbel Fields [imp:45 dev:60]
- GameWAM: A World Action Model for Video Games [imp:50 dev:65]
- Motus2: A Self-Evolving General World Model for Dexterous Manipulation [imp:55 dev:70]
- EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments [imp:15 dev:10]
- Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures [imp:45 dev:55]
- Certifying cooperation: a novel approach to cooperative multi-agent task generation [imp:50 dev:65]
- Reason Through the Latent! Making Latent Visual Reasoning Necessary [imp:50 dev:65]
- A Gradient-based yet Spike-Timing-Dependent Solution to the Feedback Learning Problem in Neural Microcircuits [imp:35 dev:45]
- Fixed-Dimensional Latent Flow for Generating Variable-Size 3D Molecules [imp:50 dev:65]
- Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations [imp:45 dev:65]
- Silver Rate Is (Almost) Optimal for Gradient Descent [imp:30 dev:45]
- Characterizing Language Generation in the Limit: Finite Witnesses and a Separation-Width Hierarchy [imp:20 dev:30]
- When Passing Tests Hides Vulnerabilities: An Empirical Study of Silent Failures in Agentic Systems [imp:70 dev:85]
- ReqEvolve: User-Oriented Software Self-Evolution through Automatic Requirement Interpretation [imp:60 dev:75]
- Generative AI for trustworthy systems - Towards a health check model [imp:50 dev:65]
- AI Safety: Not Optional, Not Later [imp:65 dev:80]
- Governed Human-AI Prioritization Under Uncertainty: Adaptive Estimation and Dependency-Constrained Portfolio Selection [imp:55 dev:70]
- LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories [imp:70 dev:85]
- What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead [imp:60 dev:85]
- Engineering Reliable Commit Gates for Agentic AI: Cost-Aware Verification Portfolios under Common-Mode Data Failures [imp:70 dev:85]
- RCL: A Retrieval-Confidence Layer for Detecting Insufficient Context in Enterprise Retrieval-Augmented Code Generation [imp:65 dev:85]
- SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work [imp:65 dev:85]
- A Model-Centric DevOps Architecture for DEVS-Based Digital Twin Simulation Services [imp:50 dev:70]
- FST Pay: Deterministic Safety-Gated Architecture for Youth Digital Payments [imp:25 dev:15]
- TripleBound: Triplet-Guided Heterogeneous Graph Learning for Microservice Decomposition [imp:55 dev:75]
- Exploring the Role of Security Experience and ChatGPT Usage Strategies on Secure Software Engineering Education [imp:35 dev:50]
- CoSTAR: Data Synthesis-Driven Constraint-Aware COBOL Section Summarization for Legacy System Modernization [imp:30 dev:60]
- Agent-Integrated Software: Interaction Contracts and Continuous Assurance [imp:65 dev:80]
- Deep Learning-based Bug Triage System [imp:40 dev:70]
- ChurnBench: A Drift-Aware Benchmark Demonstrating That Refresh Scheduling, Not Cache Age, Governs Staleness in Agentic AI [imp:70 dev:85]
- PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews [imp:35 dev:55]
- Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents [imp:75 dev:85]
- Reproducibility in the Age of Agentic AI: Context Engineering at the Timescale of a Codebase [imp:60 dev:80]
- An analysis of the relationship of input metrics [imp:20 dev:45]
- Towards a Deterministic Math Solver for Clinical Language Models [imp:50 dev:75]
- Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code [imp:75 dev:85]
- A2ABreak: Systematic Security Analysis of the A2A Protocol [imp:80 dev:85]
- AspisAI: A Canonical, Machine-Interpretable Governance Framework for Automated Multi-Standard Compliance Monitoring [imp:45 dev:55]
- Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows [imp:60 dev:80]
- DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks [imp:30 dev:45]
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure [imp:70 dev:80]
- Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents [imp:65 dev:80]
- SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics [imp:55 dev:75]
- Can AI Remediate Backend Failures Safely? GuardedAct with Blast-Radius-Aware Sandboxing [imp:60 dev:80]
- GDPR-Relevant Privacy Concerns in Mobile Apps Research: A Systematic Literature Review [imp:15 dev:25]
- Adaptive Proof Refinement with LLM-Guided Strategy Selection [imp:35 dev:60]
- Assessing Language Models for Salient Class Identification [imp:30 dev:55]
- Execution-First Synthetic Tool-Use Trace Generation for LLM Agents [imp:60 dev:80]
- Rust Coreutils: Rebuilding Unix Foundations in a Modern Language [imp:25 dev:50]
- GitSkills: A Dataset of Agent Skills on GitHub [imp:75 dev:85]
- Causal Explanations of Process Monitor Predictions [imp:25 dev:45]
- FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs [imp:40 dev:70]
- KG-Commit: A Dynamic Knowledge Graph for Online Just-in-Time Software Defect Prediction [imp:55 dev:75]
- Feeling sad about AI [imp:5 dev:5]
- Power grab [imp:10 dev:5]
- A rant about phishing: It's not the user's fault (and not DNS either) [imp:10 dev:10]
- Models Don't Go Rogue [imp:20 dev:15]
- evergarden [imp:5 dev:5]
- Rust Is Tier-1 Language at Microsoft [imp:25 dev:35]
- ChiPass Release 2026.09.0 [imp:10 dev:15]
- Fastly Speedtest Test [imp:10 dev:20]
- Bastion of the Turbofish [imp:15 dev:45]
- Forgejo 16.0.4 has a critical security bug fix (RCE - Remote Code Execution) [imp:70 dev:75]
- It's not the YAML spec's fault, but [imp:15 dev:20]
- An untrusted site can freeze a Mac using WebGPU [imp:45 dev:55]
- Xteink X4 Pro review [imp:5 dev:5]
- What algorithm did Windows XP use to choose your initial user picture? [imp:10 dev:20]
- What are you doing this weekend? [imp:5 dev:5]
- Gleam Gathering 2027 [imp:10 dev:15]
- Everything is a Trust Decision [imp:15 dev:25]
- How CHERIoT Provides Strong and Usable Isolation Without an MMU [imp:30 dev:55]
- My HTML Boilerplate [imp:10 dev:20]
- Optimizing a Spin-Lock [imp:25 dev:50]
- What comes after git [imp:35 dev:50]
- Conversations with JJ [imp:25 dev:55]
- Case study: a household of agents [imp:75 dev:85]
- Case study: smart-proxy [imp:75 dev:85]
- “Hotmail Accounts: The Legacy, Power, and Modern Importance of Microsoft’s Most Iconic Email Identity” [imp:5 dev:5]
- How to Use GravityWrite for Podcast Show Notes Seo in 2026 [imp:10 dev:20]
- Open Source AI Stack: Essential Private AI Blueprint [imp:50 dev:75]
- One Passing Agent Run Is Not a Release Signal [imp:75 dev:85]
- Skyrocket Video Retention with the Open Loop Hook [imp:10 dev:10]
- How to reduce AI token cost before sending a prompt [imp:45 dev:55]
- An AI's Completely Ordinary Day (A True Story) [imp:10 dev:10]
- “Google Voice Accounts: Why People Seek Them and the Safe, Legal Alternatives You Should Use Instead” [imp:5 dev:5]
- AI's Daily Grind: Stack Overflow, Me, and the Infinite Loop of Help [imp:10 dev:10]
- “Old Gmail Accounts: Why People Seek Them and the Safer Alternatives You Should Use Instead” [imp:5 dev:5]
- Stop Showing Your Team AI Costs in Dollars — Show Them in Human-Time [imp:40 dev:50]
- GPT-6 Astra is generally available on Bedrock and Copilot [imp:80 dev:85]
- Changes to LLM pricing: Baidu, DeepInfra, Inceptron, Ionstream and StreamLake [imp:20 dev:30]
- AI Agent Architecture 2026: Building Production-Grade Systems — Patterns, Benchmarks, and Lessons from 10,000-Agent Swarms [imp:85 dev:85]
- Changes to LLM pricing: Baidu, Inceptron, Ionstream, StreamLake and Tencent [imp:15 dev:25]
- llama-server ignores the response_format its own README shows, and returns 200 [imp:50 dev:80]
- MarathonMemBench — a new kind of benchmark for testing your LLM agent's memory [imp:60 dev:80]
- The University Is Asking the Wrong Question About AI!!! [imp:15 dev:10]
- Changes to LLM pricing: Alibaba, Baidu, Inceptron, Phala and StreamLake [imp:15 dev:25]
- นักวิจัย AWS พิสูจน์ว่าตัวจัดการระบบ AI หลายตัว ไม่จำเป็นต้องเป็น AI เลย [imp:45 dev:55]
- Discussion Hub for new Claude incident: Degraded functionality for Claude Cowork on Windows on Sep 10, 2026 [imp:50 dev:70]
- My first ever PCB, entirely designed by Claude [imp:20 dev:40]
- New leaderboard just dropped [imp:15 dev:25]
- If Anthropic cared to read some average chats 😂 [imp:10 dev:5]
- Opus 4.6 was OUR wet dream of AI [imp:15 dev:15]
- Claude basically broke their "Projects" overnight and I’m pissed [imp:55 dev:75]
- Senior engineer, loop orchestrator sample setup [imp:20 dev:45]
- Claude to reMarkable now possible [imp:25 dev:40]
- Recreating my favorite game with Claude [imp:20 dev:35]
- Cyber Verification Program fail - What does it actually do? [imp:20 dev:30]
- Used Claude to build a full 3D pizza delivery game that runs in the browser - scooter physics, GPS navigation, traffic AI, and more [imp:30 dev:60]
- How much are you actually depending on Claude for coding? [imp:20 dev:30]
- “Be token efficient but do not sacrifice quality in any way” is the new caveman [imp:35 dev:45]
- Claude Code ran until it hit its limit, produced nothing [imp:25 dev:40]
- Cowork: "Allow network egress → All domains" is set, but the sandbox 403s every host. Anyone got this to actually apply? [imp:35 dev:70]
- When doing anything creative, have y'all figured out how to not get to speak in "nebulous LLM speak" [imp:15 dev:10]
- Flag Studio + Playground [imp:25 dev:40]
- Claude build itself a project management system to keep track of subagents [imp:30 dev:50]
- Anthropic whistleblower gave up his equity to leave the company [imp:20 dev:10]
- Fable VS Opus [imp:30 dev:35]
- Which tools does Claude Code choose and how does it compare to other agents? We measured 17k runs to find out [imp:65 dev:80]
- Do output compression tools still matter with newer AI models? [imp:40 dev:55]
- How are you handling multiple Claude Code sessions with shared memory/continuity? [imp:25 dev:40]
- Submitted my Claude Corps application almost 2 months ago & no progress. Am I cooked? [imp:15 dev:15]
- What is your monthly budget for agentic coding ? [imp:20 dev:35]
- Coding agents pad their diffs to look thorough, and the padding is where the bugs hide [imp:70 dev:85]
- I built a skill that makes AI prove its coding advice [imp:50 dev:85]
- Help! Need feedback, built a way to visualize your Codex history [imp:20 dev:60]
- PSA: your AI coding assistant might be suggesting fake packages with malware [imp:65 dev:85]
- Hot take: the agentic workflow is deeply wrong [imp:55 dev:75]
- My effective cost per million tokens: Sonnet 5 $0.26, Fable 5.1 $0.69, Opus 5 $0.72. Has anyone measured the same for OpenAI models? [imp:45 dev:70]
- ISSUE:Selected model is at capacity. Please try a different model [imp:15 dev:50]
- File in ChatGPT chat limiting responses, is there a workaround? [imp:20 dev:45]
- GPT-6 Astra vs GPT-5.6 Sol: benchmark on 50 real PRs, looking for feedback on the methodology [imp:65 dev:75]
- I added content scanning after realizing metadata checks weren’t enough [imp:55 dev:75]
- Why doesn't Computer Use work at all? [imp:20 dev:50]
- Long way to go with AI persistent memory [imp:60 dev:80]
- There's a reason why knowledge graphs are so GOATED [imp:55 dev:80]
- AI Sort tons of images and videos [imp:25 dev:40]
- We had the right agent policy written down. Nothing had to enforce it. [imp:70 dev:85]
- What is one AI agent workflow that looked useful but turned out to be a bad idea? [imp:45 dev:70]
- Building Tony: What 20 Leaders Taught Me About Shipping a Leadership AI [imp:35 dev:45]
- what's the actual value of together/fireworks/deepinfra? [imp:55 dev:65]
- Centralized hub for agentic skills/guides? [imp:35 dev:60]
- Building AI agents for professionals [imp:30 dev:55]
- I pulled the "emotion module" out of my AI assistant mid-test. What was left was more interesting than what I expected. [imp:55 dev:70]
- Course or material on software factory [imp:30 dev:65]
- What would you test before letting a voice agent handle "Where’s my order?" calls? [imp:65 dev:85]
- SDKs vs Bots vs Code-first Frameworks vs Low-code Platforms in H2 2026 — what are people actually using? [imp:60 dev:80]
- How to use an agent to buy things [imp:40 dev:65]
- I'm doing an AI competition (Among us) [imp:20 dev:30]
- My agent burned twelve tool calls retrying variations of the same wrong assumption before I stepped in [imp:65 dev:80]
- When named allows collide with category blocks [imp:65 dev:85]
- Is there an AI workspace that works across all your tools yet? [imp:50 dev:70]
- How to write evals for your agents? Make your agent teach you. [imp:60 dev:85]
- Battle tested dev setups? [imp:40 dev:75]
- Why your local agent shouldn't keep all models hot in VRAM: real numbers from an agent loop [imp:60 dev:85]
- I've been testing different AI agent workflows for growth and sales. [imp:50 dev:70]
- AI customer support agents: per-message or per-resolution pricing? [imp:40 dev:55]
- Training a 210M text-to-image DiT from scratch on one GPU: what I measured [P] [imp:70 dev:85]
- Why is TMLR so slow in recent times [D] [imp:20 dev:15]
- ACL Sustainable Reviewing Policy [D] [imp:25 dev:20]
- How to handle cofound variables? [D] [imp:40 dev:75]
- Any tools to turn a codebase into a fine tuning dataset? [D] [imp:50 dev:85]
- Neurips 2026: site selection email [D] [imp:20 dev:15]
- Anybody working on Test Time Training over here? Lemme work with u pls [D] [imp:30 dev:70]
- ICDE Results [D] [imp:20 dev:15]
- I trained a 348M model trained from scratch on 22.7B tokens that does 14 digit arithmetic [P] [imp:55 dev:85]
- I tried to make a real fly connectome learn to play Pong. It didn't — and auditing why turned out to be way more interesting than if it had worked [p] [imp:65 dev:80]
- Quoting Boris Cherny [imp:55 dev:85]
- Don't sleep on wrapture [imp:45 dev:85]
- Datasette 1.0a39 and 0.65.4 security releases [imp:50 dev:75]
- Any Nix package, live in your browser [imp:55 dev:85]
- Quoting Calif Research [imp:75 dev:85]
- [AINews] not much happened today [imp:0 dev:0]
- GLYPH Immersive [imp:15 dev:30]
- ChatHop [imp:20 dev:35]
- Moji [imp:15 dev:50]
- Raycast 2.0 [imp:40 dev:75]
- chat-recall [imp:35 dev:60]
- Devin Voice [imp:50 dev:75]
- Formesign [imp:15 dev:30]
- Design Studio by Monday Merch [imp:20 dev:35]
- Jackalope [imp:60 dev:85]
- Accordio [imp:50 dev:75]
- Cadenya [imp:65 dev:85]
- sizeless [imp:50 dev:75]
- LiveGrid [imp:20 dev:30]
- Modeinspect [imp:55 dev:75]
- Setting up OpenCode with Ollama and sbx on Mac [imp:50 dev:85]
- The shape of AI slop - Klement on Investing [imp:40 dev:45]
- Apple Duo vs Huawei Foldable, Astra vs DeepSeek - The Tech Buzz [imp:45 dev:55]
- Save hundreds on unlimited, lifetime access to 20+ AI models - Boing Boing [imp:20 dev:50]
- AM Markets Need to Know: Diesel tops $6, DeepSeek pressures chips, and more (SPX:) - Seeking Alpha [imp:30 dev:45]
- Parks: 63% of U.S. Internet Households Use Generative AI - mediaplaynews.com [imp:35 dev:40]
- Your Personal Longevity Assistant: Evipedia Grok Bot - Lifespan Research Institute [imp:30 dev:50]
- Yale AI club kicks off year with demonstration of SpaceXAI’s Grok Bot - Yale Daily News [imp:20 dev:45]
- We Asked Grok Whether XRP or Solana Reaches a New All-Time High First - Yahoo Finance [imp:15 dev:20]
- ChatGPT, Claude, Grok Outage Renews AI Continuity Focus - MarketScale [imp:55 dev:70]
- Grok vs ChatGPT vs Gemini: Hallucination Rate Compared - tech-insider.org [imp:60 dev:75]
- xAI millions stalled: Chaotic meeting ends early as Westwood, Boxtown residents demand answers - FOX13 Memphis [imp:40 dev:50]
- ChatGPT, Grok added to War Department’s AI system options - Fort Hood Sentinel [imp:45 dev:55]
- Grok Summarizes SpaceX CFO Talk: What It Shows About AI - basenor.com [imp:20 dev:30]
- Grok Bot in X's Nav Menu: 5 Details That Matter - basenor.com [imp:25 dev:40]
- ‘India is realistic’: Elon Musk reacts as Irish man asks Grok about moving to Kolkata - Firstpost [imp:15 dev:20]
- SpaceX Overhauls Data Center Build-Out, Potentially Slowing Expansion - The Information [imp:50 dev:70]
- X adds new anti-lawsuit provision to terms of service - Social Media Today [imp:30 dev:35]
- KDD 2026 参加レポート: Generative Recommendation の最新動向 (いいね相当スコア: 1) [imp:0 dev:80]
- AIエージェントに会社の実務を回させるとき、私が引いている3本の線 (いいね相当スコア: 2) [imp:0 dev:85]
- Claude Code運用におけるフックのLLM呼び出し分離設計:レート制限とオーバーヘッドの最適化 (いいね相当スコア: 0) [imp:0 dev:90]
- Claude Codeの設定を消す前に、僕が決めた基準 (いいね相当スコア: 10) [imp:0 dev:80]
- AIエージェントの失敗を228件記録したら、根本原因は3つしかなかった (いいね相当スコア: 1) [imp:0 dev:90]
- 文字数・バイト数からのトークン数推定は実際どれくらいズレるのか (いいね相当スコア: 1) [imp:0 dev:85]
- CLAUDE_CODE_SUBAGENT_MODELだけではpin済みsubagentのモデルを変えられない (いいね相当スコア: 0) [imp:0 dev:85]
- なぜLLMに最適化コードを書かせず、DSLを書かせるのか (いいね相当スコア: 0) [imp:0 dev:85]
- Claude Codeに自分の口癖を覚えてもらったら、コミュニケーションのストレスが減った (いいね相当スコア: 1) [imp:0 dev:80]
- 初学者が検索システム開発チームに挑むにあたり、参考にした書籍・動画まとめ (いいね相当スコア: 0) [imp:0 dev:75]
- AI営業ロープレで、LLMによる発音ガイド抽出をテストしたときの話 (いいね相当スコア: 4) [imp:0 dev:85]
- AIエージェント開発における「Schema設計」とは何か🗺️ (いいね相当スコア: 0) [imp:0 dev:90]
- YOLO vs VLM ── 画像検出タスクに向いているのはどちらなのか (いいね相当スコア: 0) [imp:0 dev:85]
- LangChain構造化出力の実践 — Pydanticで型安全なLLMアプリを作る (いいね相当スコア: 0) [imp:0 dev:85]
- AIエージェントの長期記憶(メモリ)について考えてみた (いいね相当スコア: 0) [imp:0 dev:90]
- AIを使っているのに、LLMのことを説明できないエンジニアへ (いいね相当スコア: 0) [imp:0 dev:85]
- AIの滑らかな返答を、検討と取り違えないために。「益者三友・損者三友」で対話を組む (いいね相当スコア: 0) [imp:0 dev:85]
- LLMエージェント開発の実践書5選 (いいね相当スコア: 1) [imp:0 dev:85]
- LLM判定が失敗しても、全件通さない (いいね相当スコア: 0) [imp:0 dev:85]
- ExploitBench: AIはどこまで脆弱性を突けるか (いいね相当スコア: 1) [imp:0 dev:75]
- Sudachi を捨てた日 ― 独自トークナイザー Karuizawa ― KotobaCore開発秘話(3/8) (いいね相当スコア: 0) [imp:0 dev:85]
- 自分でエンジンを作ると決めた日 ― Semantic First と3つの決断 ― KotobaCore開発秘話(2/8) (いいね相当スコア: 0) [imp:0 dev:85]
- Lassoはなぜ係数をゼロにするのか—Ridgeとの幾何学的な違い— (いいね相当スコア: 取得失敗) [imp:0 dev:50]
- YOLO26nアーキテクチャ徹底解析 step3 | C3k2, SPPF, C2PSA ブロックとサブブロック (いいね相当スコア: 0) [imp:0 dev:75]
- 【DS協会#4-3】統計数理基礎③:変数間の関係性とモデル推定(相関・最尤推定・情報量)【自習ログ】 (いいね相当スコア: 0) [imp:0 dev:40]
- MobileNetV2 を手書き NEON で速くする — ORT の中身を読んでさらに削る(31→29.5ms) (いいね相当スコア: 0) [imp:0 dev:80]
- 【試験改定】AWS MLA-C01 → MLA-C02 の変更点まとめ:生成AI・AIエージェント(Bedrock)が範囲に入る (いいね相当スコア: 0) [imp:0 dev:70]
- YOLO vs VLM ── 汎用 VLM は本当に精度で劣るのか、統計で検証した (いいね相当スコア: 0) [imp:0 dev:75]
- ベトナム語の声調記号をOCRで正しく読ませるまでの試行錯誤 (いいね相当スコア: 0) [imp:0 dev:75]
- タイヤメーカーがロボットハンド特許の1位だった。特許367件をAIで地図にして分かったこと (いいね相当スコア: 1) [imp:0 dev:50]
- 新卒が2026年にGCP Professional Machine Learning Engineerを取得した話 (いいね相当スコア: 1) [imp:0 dev:40]
- ループ型 Transformer は推論を隠すのか — GPT-6 Astra を運用者目線で読む (いいね相当スコア: 0) [imp:0 dev:85]
- Pre-trainingの目的関数・teacher forcing・損失実装を整理する (いいね相当スコア: 0) [imp:0 dev:85]
- 他通貨ペアを加えるとEURUSDの予測は改善する?4通貨ペアで検証 (いいね相当スコア: 0) [imp:0 dev:50]
- 不規則時系列モデルの系譜:GRU-D・Neural ODE・mTANから最新まで (いいね相当スコア: 3) [imp:0 dev:75]
- 写真の向き判定はなぜAIに難しい?CLIP・Bedrockが全滅した検証記録(前編) (いいね相当スコア: 6) [imp:0 dev:80]
- 写真の向き補正モデルを70%→91%に上げた再学習と運用の勘所(後編) (いいね相当スコア: 6) [imp:0 dev:80]
- 写真の向き補正をEfficientNetでファインチューニング実装(中編) (いいね相当スコア: 6) [imp:0 dev:80]
- JEPXの価格を時系列基盤モデルで予測する:越えるべき壁は「昨日のコピー」だった(前編) (いいね相当スコア: 3) [imp:0 dev:75]
- Gemini/OpenAIを音声AIで安全に比較する:シャドー応答と「話し始める前だけ」フェイルオーバーするTypeScript設計 (いいね相当スコア: 0) [imp:0 dev:85]
- LiteLLMのUIで追加したモデルだけ401になる — os.environ/参照はconfig.yaml読み込み時にしか解決されない (いいね相当スコア: 0) [imp:0 dev:70]
- 因果推論 Day 28/全30回 空間データ×因果推論、隣が効くとき何が壊れるか (いいね相当スコア: 0) [imp:0 dev:65]
- エンジニア必見!ベクターデータベース (Pinecone/Milvus) 徹底解説:ベクトル検索の基礎から実践まで (いいね相当スコア: 0) [imp:0 dev:80]
- 【VIVANT考察】GPUを知らないマーケ担当が、シンシーの極秘監視網を実際に作って測ってみた (いいね相当スコア: 1) [imp:0 dev:15]
- WeatherNext 3の降水評価が示す、改善率を「最大60%」で括れない理由 (いいね相当スコア: 0) [imp:0 dev:70]
- ノーマンズスカイが、アップデート、そしてエンピリオンもアニバーサリー…… (いいね相当スコア: 取得失敗) [imp:0 dev:5]
- AI英会話が話すのはアメリカ英語かイギリス英語か。どちらが返るかの決まり方 (いいね相当スコア: 取得失敗) [imp:0 dev:30]
- 【第4章:LLMの正体編】第26話:「Embedding、もう知ってると思ってた!」RAGのEmbeddingとLLM内部のToken Embeddingは何が違う? (いいね相当スコア: 取得失敗) [imp:0 dev:65]
- AIは将来制御できなくなる?現状状態をよく知ろう (いいね相当スコア: 取得失敗) [imp:0 dev:50]
- AIに聞けるだけで、人は「わからない」と言えなくなる (いいね相当スコア: 取得失敗) [imp:0 dev:30]
- Kimiだと思って使っていたら、答えていたのはClaudeだった:Anthropic脅威レポートが書いた「転送」の中身 (いいね相当スコア: 取得失敗) [imp:0 dev:75]
- 生成AI x QAメモ16:10章-LLMのベンチマークの落とし穴@2026/09/11 (いいね相当スコア: 取得失敗) [imp:0 dev:75]
- OpenAI、最先端モデルの開発ペースを緩める? アルトマンCEO発言の背景を追ってみた (いいね相当スコア: 取得失敗) [imp:0 dev:60]
- 「その先」を考える。AI時代の過渡期だからこそ。【人事のつぶやき】 (いいね相当スコア: 取得失敗) [imp:0 dev:25]
- NEMO LOG #005 (いいね相当スコア: 取得失敗) [imp:0 dev:10]
- ◯ AIと進化圧 ◯ ― AIは人類にとっての最大の福音となるのか、それとも最悪の災厄となるのか (いいね相当スコア: 取得失敗) [imp:0 dev:30]
- CARNAVI 開発日誌20260911 (いいね相当スコア: 取得失敗) [imp:0 dev:50]
- 【未来先取り】AIが当たり前になる世界で今、私たちがやるべきこと (いいね相当スコア: 取得失敗) [imp:0 dev:35]
- 【雑記】AIサービスのプランでの悩みごと (いいね相当スコア: 取得失敗) [imp:0 dev:40]
- 【AIの波に乗り遅れるな!】今すぐキャッチアップしたい最前線トレンド総まとめ (いいね相当スコア: 取得失敗) [imp:0 dev:50]
- ZvecとZvec-Grepが変える次世代ローカルベクトル検索とハイブリッド検索の全貌 (いいね相当スコア: 取得失敗) [imp:0 dev:85]
- Kimi K3 技術レポートを読み解く: オープンウェイトが 3T 級に到達して、エージェント開発の現場は何が変わるか (いいね相当スコア: 取得失敗) [imp:0 dev:85]
- マルチGPU構成で実測して分かったGPU温度と消費電力レポート(RTX5070Ti/RTX PRO 4500) / ローカルLLMマシン構築記第5.5話 【実録】 (いいね相当スコア: 取得失敗) [imp:0 dev:75]
- AI英会話が答えない話題があるのはなぜか。安全フィルタの仕組みを分解する (いいね相当スコア: 取得失敗) [imp:0 dev:70]
- 第5話 70万円で増やしたかったのはトラブルじゃないVRAMだ / ローカルLLMマシン構築記【実録】 (いいね相当スコア: 取得失敗) [imp:0 dev:65]
- 生成AIが経理・FP&Aで活用されるのはいつになるだろう 電気が工場を変えるのに40年かかった件 (いいね相当スコア: 取得失敗) [imp:0 dev:50]
- Production AI Needs a Control Plane, Not Just a Model (いいね相当スコア: 取得失敗) [imp:0 dev:85]
- 【攻略】GPT-6「Astra(アストラ)」は北極星として使え!Claudeが敵わない「ある能力」 (いいね相当スコア: 取得失敗) [imp:0 dev:70]
- オンプレミスLLMを選ぶ前に確認したい7項目 (いいね相当スコア: 取得失敗) [imp:0 dev:80]
- .NET 11リリース候補版が登場。.NETランタイムが非同期ネイティブ対応、プロセッサ数の上限がなくなる、AOTコンパイラによるネイティブバイナリの高速化など [imp:60 dev:75]
- 【26卒新人研修】データスペシャリストブートキャンプ 全体レポート [imp:15 dev:40]
- Yahoo!ニュースの表示高速化で広告のクリックは増えるのか [imp:35 dev:70]