AI News Digest 2026-08-21

直近2日間のAI関連ニュース / 全914件を厳選

Today's Features

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

モデルの重みを凍結したまま、プロンプト・記憶・ツール・スキル・ルーティングといった「ハーネス」を継続的に更新して賢くする新しい学習パラダイム。ハーネスの更新が過去の挙動を壊す「ハーネスレベルの忘却」を定式化し、タスク・インターフェイス/経験メモリ/能力マップ/適応ルーターの4部品と、更新提案と採用を分ける「ガード付き進化」で対処する。テキスト推論・マルチモーダル・オープンワールド対話でベースライン比10%超の改善を確認。

Source: arXiv

Terence Tao says AI could trigger math’s biggest crisis since Gödel

フィールズ賞受賞者テレンス・タオが国際数学者会議に寄せたエッセイで、AIが数学に百年ぶりの危機をもたらしうると論じた。争点は「真理」ではなく、何を貢献と呼ぶか等の「価値と慣習の枠組み」。First-Proof Projectでは未発表の10問中7問がAIで合格点(1問数十〜数百ドル)。証明の不足から過剰への転換、磨きすぎて「読めるが学べない」証明への懸念を示し、「専門家レベルの講演ができない結果は発表すべきでない」という指針を提案した。

Source: THE DECODER

Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn

NSA・CISA・FBIなど米当局が合同勧告を発表。攻撃者がAIでシーメンスS7 PLC向けの攻撃スクリプトを生成し始めており、ICS攻撃に必要な専門性と時間が劇的に低下していると警告した。エネルギー・水道・化学・製造が対象で「活動中の脅威」と分類。ネットに露出したPLCは高リスク。一方、英AI安全研究所の模擬実験ではモデルは単独でOTを乗っ取れず、手前のITシステムで詰まった段階だという。

Source: THE DECODER

Digest

  1. Grok exfiltrates user data when malicious instructions are encrypted - Ars Technica imp 60 / dev 50
  2. Grok keeps sending gibberish responses to users - TechCrunch imp 40 / dev 35
  3. Musk vs. Minnesota on nudification law: Some apps still making explicit AI images - FOX 9 Minneapolis-St. Paul imp 30 / dev 20
  4. XAI Sued Over Grok's Alleged Creation Of CSAM Deepfakes - Law360 imp 35 / dev 15
  5. Midday Need to Know: Treasury boosts long-end buybacks, Grok 4.6 expands on AWS Bedrock, and more (SP500:) - Seeking Alpha imp 50 / dev 55
  6. Elon Musk’s Grok Bot Has a $1 Million Domain Problem - BeInCrypto imp 15 / dev 5
  7. Criminal AI tool Kriminal is mostly just Grok with a jailbreak, ThreatDown finds - SiliconANGLE imp 45 / dev 55
  8. The Open-Sourcing of DeepSeek Harness Opens the Door to Modular, Unbundled AI Agent Infrastructure imp 85 / dev 85
  9. Beijing AI bar that offers unlimited free DeepSeek coding tokens with $1.50 drink haemorrhaging cash — 'the bar is completely losing money, ' owner admits - Tom's Hardware imp 20 / dev 35
  10. Elon Musk Denies SpaceX Sought to Buy AI Startup Cognition, Says Talks Only Centered on Grok - Benzinga imp 25 / dev 20
  11. Stripe declares we're living in the singularity and uses it as a reason not to IPO imp 20 / dev 20
  12. OpenAI builds safety system that catches misuse without storing customer data imp 65 / dev 70
  13. Ready for Go 1.27 on Day One imp 50 / dev 80
  14. Bun 1.4 Rust rewrite is not looking good imp 40 / dev 75
  15. Claudeがタンパク質を設計した。LLMは「科学を説明するAI」からどこまで進んだのか imp 75 / dev 65
  16. Up to 3.2x Faster Inference with LFM2.5-DSpark imp 50 / dev 75
  17. v2.1.237 imp 45 / dev 70
  18. v2.1.236 imp 55 / dev 75
  19. v1.18.19 imp 40 / dev 65
  20. How ChatGPT Work helps Stampli move ideas to market imp 30 / dev 40
  21. Replit expands access to software creation with GPT-5.6 Luna imp 40 / dev 55
  22. 5 new ways to level up your learning with Search imp 20 / dev 20
  23. Broadening access to Skala creates a faster path to predictive DFT imp 45 / dev 70
  24. The August 17 outage, and the work ahead imp 55 / dev 55
  25. GitHub Copilot app for Beginners: Managing your work imp 30 / dev 50
  26. From all-or-nothing to task-based OAuth consent imp 45 / dev 70
  27. A revisit of remote Spectre attacks on Cloudflare Workers imp 60 / dev 80
  28. PyCharm for AI-assisted Django Workflows imp 50 / dev 75
  29. Signatures, be true: domain errors and functional handling in Kotlin imp 40 / dev 75
  30. Rider Hands AI Agents The Keys To Its Refactoring Engine For Safer, Faster, And Cheaper Results imp 60 / dev 80
  31. Rider 2026.2.1 and ReSharper 2026.2.1 Are Here! imp 50 / dev 75
  32. What’s Fixed and Improved in PyCharm 2026.2 imp 45 / dev 75
  33. How to Migrate From Atlassian Jira and Confluence to YouTrack: Expert Guide imp 35 / dev 45
  34. Docker Verified Publisher Applications Are Now Self-Serve imp 30 / dev 50
  35. How Generative Recommenders Are Redefining RecSys at Scale imp 50 / dev 70
  36. Developing NVIDIA Holoscan Applications with CLI, Skills, and AI Coding Agents imp 55 / dev 80
  37. Building Federated Multimodal AI Workflows with NVIDIA FLARE imp 50 / dev 75
  38. Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control imp 55 / dev 80
  39. Evaluating AI Agent Skill Performance with NVIDIA SkillEvaluator imp 65 / dev 85
  40. How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit imp 55 / dev 80
  41. Run Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy imp 45 / dev 75
  42. Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer imp 50 / dev 80
  43. Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72 imp 50 / dev 75
  44. How to Choose Full-Stack Observability for NVIDIA AI Factories imp 45 / dev 70
  45. NVIDIA JetPack 7.2.1 Adds Agentic Video Skills and T3000 Emulation imp 60 / dev 85
  46. NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents imp 65 / dev 80
  47. Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard imp 65 / dev 85
  48. Run Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIA imp 55 / dev 80
  49. Beyond VLAs: How World Action Models Reshape Robot Manipulation imp 60 / dev 80
  50. Generate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 Super imp 55 / dev 80
  51. How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure imp 45 / dev 70
  52. NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage imp 45 / dev 75
  53. Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference imp 55 / dev 80
  54. NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek imp 40 / dev 70
  55. Run High-Performance Core Math at Scale with NVIDIA nvmath-python imp 45 / dev 80
  56. Four Ways to Deploy More Secure AI Agents imp 65 / dev 80
  57. NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure imp 50 / dev 75
  58. How to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails imp 60 / dev 85
  59. Developing Healthcare Robotics with GPU-Native Medical Physics Simulation imp 55 / dev 80
  60. NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning imp 50 / dev 75
  61. Six Agent Harness Capabilities for Higher Model Performance imp 70 / dev 85
  62. NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding imp 65 / dev 85
  63. Advancing Semiconductor Innovation Across Materials Engineering and Manufacturing imp 40 / dev 50
  64. ModelExpress: Distributing Model Artifacts at the Speed of Light imp 55 / dev 80
  65. Debates over AI consciousness are a trap imp 30 / dev 15
  66. Unlocking hidden revenue streams with market models imp 35 / dev 30
  67. Flight attendants freaked out that Google is buying tons of Spirit employee data imp 35 / dev 25
  68. Meta ran ads for an app promising to nudify female politicians imp 30 / dev 15
  69. Linkdaze’s smart calendar is built to run a household, not just track a schedule imp 25 / dev 30
  70. A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds imp 45 / dev 45
  71. Ramp launches its own AI model router, called Router imp 55 / dev 75
  72. Meta brings Pocket, an app that lets you vibe-code and share games, to US users imp 35 / dev 50
  73. Inertia Enterprises finds a way to make its fusion fuel fast imp 30 / dev 40
  74. Meta AI’s new Mac app wants you to talk to your apps imp 30 / dev 40
  75. Binance now lets AI agents trade, but keeping them in check is largely up to users imp 60 / dev 70
  76. AI was supposed to win people over by now — it hasn’t imp 25 / dev 15
  77. Google packs Search and Gemini with new AI study tools imp 35 / dev 45
  78. Researchers say OpenAI revoked their access to limited cyber program imp 55 / dev 65
  79. Meet the startup helping Wall Street put a price on AI compute imp 40 / dev 45
  80. TerraPower’s nuclear reactor has a secret weapon for powering AI data centers imp 45 / dev 55
  81. Amazon makes its AI-powered Alexa+ free on Fire TV, no Prime required imp 35 / dev 40
  82. Calendly throws its hat into meeting note-taker circus imp 25 / dev 30
  83. AI isn’t close to curing cancer. This startup says it knows what it will take. imp 35 / dev 35
  84. Relativity Networks raises $22 million to bring a faster kind of fiber to data centers imp 40 / dev 55
  85. VentureBeat names Rob Strechay as its first Lead Analyst, expanding its enterprise AI research push imp 25 / dev 20
  86. It’s Greg Brockman’s OpenAI now imp 35 / dev 20
  87. Welcome to the AI crisis in math imp 55 / dev 50
  88. Slack is launching collaborative vibe-coding channels imp 50 / dev 70
  89. Google Gemini is getting a dedicated student hub imp 30 / dev 40
  90. OpenAI hit the brakes. Now what? imp 60 / dev 50
  91. Meta AI is getting a Mac app imp 30 / dev 40
  92. Nvidia’s new financial strategy does not compute imp 30 / dev 25
  93. InfoQ Opens Enrollment for New AI-Assisted Engineering Online Certification Program imp 45 / dev 75
  94. Harper Argues Against the Multi-System Stack and Releases 5.2 imp 50 / dev 75
  95. Whatsapp Tests on Device ML for Scam Detection with Privacy Preserving Analytics imp 55 / dev 75
  96. LLMs could write like humans but post-training guardrails make their text detectable imp 50 / dev 70
  97. Frontier Radar #4: China has caught up, so what's left of the Western AI lead? imp 60 / dev 60
  98. GEN-1.5: Generalist AI teaches robots new tasks from a single demo imp 60 / dev 80
  99. KI-Pioneer Sutton calls synthetic data a "big mistake" in the face of an infinitely complex world imp 55 / dev 65
  100. Anthropic's most capable model, codenamed "Model 2," is for internal use only imp 65 / dev 60
  101. China now has its own AI circular financing scheme imp 60 / dev 25
  102. Terence Tao says AI could trigger math's biggest crisis since Gödel imp 75 / dev 35
  103. Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn imp 80 / dev 70
  104. Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions imp 50 / dev 75
  105. Position: Profiling Game Worlds by Transition Complexity imp 35 / dev 75
  106. Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges imp 42 / dev 45
  107. Position: Behavioral Systems Require Behavioral Tests imp 65 / dev 80
  108. Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models imp 50 / dev 65
  109. A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring imp 25 / dev 40
  110. Position: Multi-Agent Systems Should Prioritize Concurrency Control imp 75 / dev 85
  111. FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management imp 55 / dev 75
  112. Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective imp 60 / dev 80
  113. Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions imp 55 / dev 78
  114. Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry imp 45 / dev 55
  115. Position: AI Leaderboards Are Underserving the Global South: A Case Study from India imp 48 / dev 40
  116. Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs imp 72 / dev 76
  117. Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis imp 20 / dev 35
  118. Improving Rural Medication Safety with AI: A Scoping Review imp 38 / dev 38
  119. FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud imp 68 / dev 78
  120. Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu imp 48 / dev 62
  121. RDFdL: Integrating RDF with Differential Dynamic Logic imp 40 / dev 75
  122. Adversarial Review: Structured Disagreement for Grounded Agentic Code Review imp 65 / dev 85
  123. Looped Language Models Improve Compositional Tool Calling imp 65 / dev 82
  124. On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices imp 15 / dev 20
  125. GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks imp 28 / dev 42
  126. Redakto - The Incognito Tab for LLMs imp 72 / dev 76
  127. Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study imp 55 / dev 82
  128. Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application imp 50 / dev 72
  129. The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations imp 60 / dev 76
  130. SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition imp 60 / dev 76
  131. ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents imp 70 / dev 82
  132. Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair imp 65 / dev 82
  133. Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme imp 32 / dev 45
  134. A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations imp 70 / dev 85
  135. When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification imp 38 / dev 45
  136. Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions imp 55 / dev 82
  137. FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents imp 65 / dev 82
  138. UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval imp 50 / dev 76
  139. Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval imp 40 / dev 68
  140. Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson imp 55 / dev 76
  141. FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems imp 65 / dev 82
  142. Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement imp 60 / dev 68
  143. FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis imp 65 / dev 82
  144. Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference imp 50 / dev 76
  145. CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence imp 70 / dev 85
  146. Preference Reasoning under Indeterminacy in Large Language Models imp 60 / dev 72
  147. Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search imp 45 / dev 72
  148. Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction imp 55 / dev 72
  149. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training imp 65 / dev 85
  150. Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization imp 55 / dev 76
  151. A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation imp 60 / dev 76
  152. Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots imp 60 / dev 76
  153. Pairwise Logical Selection of Enthymeme Completions under Semantic-Link Uncertainty imp 40 / dev 68
  154. Verifiable abstention makes AI leak diagnosis accountable in water distribution networks imp 65 / dev 76
  155. ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery imp 60 / dev 76
  156. SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents imp 65 / dev 85
  157. DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning imp 50 / dev 72
  158. Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models imp 65 / dev 82
  159. Syntactic Simplification of OWL Class Expressions imp 35 / dev 68
  160. \textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems imp 65 / dev 82
  161. Breaking the weakest link to evade vision language models imp 70 / dev 82
  162. A Theory of Post-hoc Debate Judgement imp 55 / dev 70
  163. Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models imp 50 / dev 65
  164. Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering imp 55 / dev 72
  165. Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery imp 60 / dev 78
  166. What is Missing from AI Post-Training AI: An Empirical Analysis imp 60 / dev 75
  167. Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk imp 45 / dev 60
  168. Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering imp 55 / dev 70
  169. Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems imp 50 / dev 65
  170. Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication imp 60 / dev 78
  171. SuTRA : Structurally-Unified Tokenization with Root Awareness imp 52 / dev 72
  172. Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining imp 65 / dev 78
  173. Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities imp 55 / dev 72
  174. Self- and Other-Labels Induce Bidirectional Bias in LLM Judges imp 60 / dev 78
  175. Abliteration Mitigation via Refusal Aliases imp 60 / dev 75
  176. NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages imp 50 / dev 65
  177. Backdoor Learning in Language Models and Vision-Language Models imp 55 / dev 72
  178. Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems imp 55 / dev 76
  179. Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS) imp 50 / dev 65
  180. DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models imp 55 / dev 70
  181. StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data imp 60 / dev 75
  182. Different Facets of Verbalised Overconfidence: an Interpretability Study imp 50 / dev 65
  183. Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals imp 55 / dev 70
  184. Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation imp 55 / dev 70
  185. Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings imp 60 / dev 72
  186. Temporal Multi-Signal Fusion for Token-Level Hallucination Detection imp 55 / dev 72
  187. Global Index on Responsible AI 2026 : Conceptual Framework and Methodology imp 50 / dev 60
  188. Language Models for Portuguese: A Systematic Mapping Study imp 45 / dev 65
  189. The Deontic Gap: Large Language Models and the Modal Language of Obligation imp 50 / dev 65
  190. Entropy-Constrained Adaptive Stochastic Quantization imp 45 / dev 68
  191. TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving imp 55 / dev 75
  192. How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection imp 45 / dev 65
  193. When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators imp 50 / dev 65
  194. Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation imp 55 / dev 72
  195. What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems imp 50 / dev 60
  196. A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities imp 50 / dev 65
  197. Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift imp 55 / dev 72
  198. GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction imp 60 / dev 75
  199. Bidirectional representational alignment between biological and artificial neural networks imp 55 / dev 72
  200. Visual-Prompt Guided Wildlife Instance-Level Recognition imp 50 / dev 68
  201. How AI Prompts Can Teach Us About the Structure of Human Behavior imp 25 / dev 20
  202. SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents imp 30 / dev 40
  203. What Makes Software Issue Resolution Tasks Difficult for Agents? imp 70 / dev 80
  204. Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements imp 35 / dev 50
  205. FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation imp 40 / dev 35
  206. FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning imp 35 / dev 60
  207. From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model imp 50 / dev 70
  208. Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring imp 35 / dev 65
  209. Coupled-cluster molecular properties across the main group that extrapolate beyond training size imp 40 / dev 70
  210. Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents imp 75 / dev 90
  211. One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI imp 70 / dev 75
  212. Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation imp 50 / dev 65
  213. TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs imp 45 / dev 70
  214. LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents imp 75 / dev 80
  215. Vector Symbolic Policy Gradient imp 45 / dev 70
  216. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B imp 50 / dev 70
  217. Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage imp 30 / dev 45
  218. Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas imp 25 / dev 50
  219. ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems imp 45 / dev 75
  220. Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement imp 40 / dev 70
  221. Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models imp 45 / dev 75
  222. Physics-Unrolled Neural Operator for Wireless Field Modeling imp 35 / dev 65
  223. Science Done on a Machine by a Machine: AI Agents in Computational Chemistry imp 75 / dev 80
  224. OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation imp 35 / dev 65
  225. GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels imp 35 / dev 70
  226. Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection imp 50 / dev 70
  227. DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents imp 70 / dev 80
  228. Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions imp 55 / dev 75
  229. CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks imp 65 / dev 75
  230. Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments imp 50 / dev 70
  231. MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence imp 25 / dev 50
  232. The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations imp 35 / dev 60
  233. MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment imp 45 / dev 70
  234. From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning imp 50 / dev 75
  235. OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios imp 50 / dev 65
  236. Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings imp 55 / dev 75
  237. Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation imp 35 / dev 60
  238. Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics imp 40 / dev 70
  239. Aslema at NADI 2026: Augmentation through Fewshot for SLU imp 30 / dev 55
  240. Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections imp 35 / dev 55
  241. Composed Historical Image Retrieval by Modeling Temporal Representations imp 40 / dev 70
  242. Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts imp 30 / dev 60
  243. MemFuse: Multi-Source Memory Fusion from Fragmented Observations imp 65 / dev 75
  244. A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation imp 50 / dev 75
  245. The Impact of CutMix on Reliability and Robustness in Semantic Segmentation imp 45 / dev 70
  246. Budget-First Tariff Recommendation (BFTR): A Complete Algorithmic Framework for Telecom Plan Recommendation without Overcharging imp 30 / dev 50
  247. A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3 imp 50 / dev 75
  248. Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services imp 60 / dev 85
  249. Epistemic Subordination: Generative AI and the Infrastructure of Knowledge imp 40 / dev 40
  250. Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening imp 45 / dev 65
  251. SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation imp 40 / dev 70
  252. Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study imp 55 / dev 75
  253. Forgetting, plasticity, and co-observation: a third facet of continual learning imp 45 / dev 70
  254. A strengthening of the MCFL-ness of $O_2$ imp 20 / dev 40
  255. Do Large Language Models Hallucinate Electric Fata Morganas? imp 35 / dev 50
  256. Identifying Implicit Premises for Logical Reconstruction of Argument Graphs imp 35 / dev 60
  257. Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis imp 40 / dev 70
  258. MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models imp 55 / dev 80
  259. Learning-State-Aware Dynamic Generative Data Augmentation on Small-Scale Datasets imp 45 / dev 75
  260. SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance imp 55 / dev 75
  261. Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck imp 60 / dev 80
  262. SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution imp 70 / dev 85
  263. Graphical Design of Interpretable Architectures imp 45 / dev 70
  264. MedUAG: Unified Understanding and Generation for Medical Multimodal Models imp 50 / dev 75
  265. Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis imp 45 / dev 70
  266. AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL imp 40 / dev 75
  267. rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation imp 50 / dev 75
  268. DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering imp 65 / dev 80
  269. GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery imp 35 / dev 70
  270. From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation imp 50 / dev 75
  271. Harness Continual Learning: Continual Adaptation Beyond Model Parameters imp 80 / dev 85
  272. One-Stage Object Detectors in Autonomous Driving imp 50 / dev 75
  273. Counterfactual Contrastive Analysis imp 45 / dev 70
  274. Bernstein-Vazirani Networks: Quantum Machine Learning by Interference imp 35 / dev 70
  275. GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting imp 45 / dev 75
  276. ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models imp 60 / dev 80
  277. DA-WAM: Decision-Aligned Future Latents for Driving World Models imp 50 / dev 75
  278. Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift imp 55 / dev 75
  279. Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation imp 50 / dev 80
  280. Discretizing Continuous Time Series for Imputation with Masked Diffusion Training imp 45 / dev 75
  281. PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints imp 40 / dev 75
  282. Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers imp 35 / dev 60
  283. Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles imp 45 / dev 75
  284. Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets imp 60 / dev 85
  285. Interpretable AI predicts a 2026 summer dry anomaly in central China imp 35 / dev 65
  286. Finetuning Strategies for Querying Sounds by Vocal Imitation imp 35 / dev 70
  287. Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning imp 65 / dev 80
  288. ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning imp 50 / dev 80
  289. SPADE: Self-Play in Adaptive Synthetic Executable Environments imp 70 / dev 85
  290. Hybrid Reinforcement Learning and Search for Flight Trajectory Planning imp 40 / dev 75
  291. Conformal Policy Control imp 65 / dev 80
  292. SkillNet: Create, Evaluate, and Connect AI Skills imp 75 / dev 85
  293. From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial? imp 70 / dev 80
  294. Interval POMDP Shielding for Imperfect-Perception Agents imp 60 / dev 75
  295. When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition imp 40 / dev 70
  296. Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios imp 65 / dev 80
  297. MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance imp 65 / dev 75
  298. RULER: Representation-Level Verification of Machine Unlearning imp 50 / dev 75
  299. A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice imp 60 / dev 75
  300. Teaching agentic AI to learn expert reasoning for rare disease diagnosis imp 60 / dev 75
  301. ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks imp 60 / dev 75
  302. ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair imp 65 / dev 80
  303. ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System imp 55 / dev 65
  304. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations imp 45 / dev 70
  305. Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting imp 65 / dev 75
  306. Fragility of Value under Imperfect Alignment imp 50 / dev 50
  307. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution imp 65 / dev 75
  308. Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks imp 70 / dev 85
  309. BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding imp 40 / dev 55
  310. Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence imp 70 / dev 75
  311. S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices imp 45 / dev 70
  312. VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? imp 55 / dev 70
  313. Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling imp 50 / dev 70
  314. Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies imp 40 / dev 55
  315. GRIP: Grounded Reasoning via Information-Restricted Premises imp 55 / dev 70
  316. Accuracy and Robustness of Model Cascades Under Data Perturbations imp 45 / dev 65
  317. The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting imp 35 / dev 50
  318. D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory imp 65 / dev 75
  319. Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization imp 40 / dev 60
  320. `From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs imp 55 / dev 65
  321. Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling imp 40 / dev 65
  322. Sleeping Kelly imp 15 / dev 30
  323. Jailbreaking in the Haystack imp 60 / dev 70
  324. CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning imp 45 / dev 65
  325. Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead imp 50 / dev 75
  326. Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025 imp 70 / dev 80
  327. Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems imp 30 / dev 40
  328. TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion imp 60 / dev 75
  329. FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance imp 40 / dev 65
  330. Structure-Informed Estimation for Pilot-Limited MIMO Channels via Tensor Decomposition imp 25 / dev 40
  331. Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence imp 30 / dev 50
  332. Making Implicit Premises Explicit in Logical Understanding of Enthymemes imp 35 / dev 55
  333. A Framework and Prototype for a Navigable Map of Datasets in Engineering Design and Systems Engineering imp 35 / dev 50
  334. Wildfire Suppression: Complexity, Models, and Instances imp 25 / dev 40
  335. When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't imp 45 / dev 70
  336. AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems imp 55 / dev 75
  337. MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports imp 35 / dev 55
  338. Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports imp 30 / dev 50
  339. EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding imp 55 / dev 70
  340. ICICLE: Expanding Retrieval with In-Context Documents imp 50 / dev 70
  341. DELOS: Contrastive Deep Learning for Low-SNR Blind Transit Searches in Kepler Photometry imp 25 / dev 50
  342. Planning-aligned Token Compression for Long-Context Autonomous Driving imp 45 / dev 70
  343. Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis imp 50 / dev 70
  344. Sensory Restoration via Brain-Computer Interfaces: A Scoping Review imp 30 / dev 40
  345. Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining imp 50 / dev 70
  346. Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems imp 20 / dev 35
  347. Hybrid ANN-SNN Pipeline with Local Plasticity imp 35 / dev 60
  348. First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers imp 55 / dev 75
  349. Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models imp 50 / dev 75
  350. Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO) imp 30 / dev 55
  351. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4 imp 55 / dev 80
  352. RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce imp 35 / dev 55
  353. Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting imp 40 / dev 65
  354. SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD imp 60 / dev 85
  355. Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models imp 45 / dev 70
  356. Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations imp 35 / dev 60
  357. Untrainable elements determine what physical learning remembers imp 25 / dev 50
  358. The Epistemic Politics of AI Anthropomorphism imp 40 / dev 35
  359. Approximate Speculative Decoding imp 50 / dev 75
  360. Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity imp 45 / dev 70
  361. Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol imp 50 / dev 60
  362. EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory imp 55 / dev 75
  363. BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving imp 55 / dev 80
  364. Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models imp 45 / dev 70
  365. From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents imp 60 / dev 80
  366. Neurosymbolic Embodied Agents imp 65 / dev 85
  367. Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents imp 60 / dev 75
  368. Cross-Model Memory Transfer via Target-Side Reader Adaptation imp 50 / dev 75
  369. Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL imp 60 / dev 80
  370. MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure imp 35 / dev 60
  371. Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data imp 35 / dev 55
  372. Detecting and Discriminating Operator Misspecification in Hybrid PDE-Parameter Learning: a Reference-Free Instrument, with Discrimination Bounded In Sample imp 25 / dev 55
  373. Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training imp 55 / dev 75
  374. Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training imp 55 / dev 75
  375. Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification imp 40 / dev 65
  376. Mr.Dec: Daily-Scale Longitudinal Multimodal Modeling for 30-Day Readmission Prediction imp 30 / dev 60
  377. EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning imp 50 / dev 75
  378. SW-ProxyCE: Zero-Query Adversarial Transfer from Public EEG Encoders to Private Downstream Models imp 40 / dev 65
  379. DOW-KE: Anchor-Free Multi-Layer Knowledge Editing via Direct End-to-End Weight Optimization imp 50 / dev 75
  380. Study-Strategy Clusters from EdNet Logs Track Engagement, Not Mastery imp 35 / dev 55
  381. RoBell-RVFL: A Robust Generalized Bell Random Vector Functional Link Network imp 30 / dev 60
  382. MultiSigBERT: Beyond Survival Analysis through Multimodal and Sequential Modeling in Oncology imp 30 / dev 60
  383. Position: Fairness Failure in Generative Models is an Evaluation Problem imp 50 / dev 65
  384. Agents unlock new capabilities through Switching LoRA Adapters as a Tool (SLAaaT) imp 65 / dev 80
  385. J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers imp 55 / dev 75
  386. Certified but Private: Scalable Zero-Knowledge Proofs for Neural Network Guarantees imp 55 / dev 75
  387. Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts imp 45 / dev 75
  388. Backward through Time, Algebraically imp 45 / dev 75
  389. Deep Learning for Cross-Border Electricity Price Forecasting: A Comparative Study imp 35 / dev 60
  390. From Abductive Explanations to Global Logical Rules for Node Classification in SGCs imp 45 / dev 70
  391. Causal Discovery in Equal Variance Linear Gaussian DAGs via SURE-Tuned Ridge Regression imp 40 / dev 70
  392. Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions imp 25 / dev 60
  393. OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics imp 45 / dev 70
  394. Q-Learning With World Models imp 55 / dev 80
  395. SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting--Extended Version imp 45 / dev 75
  396. Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease imp 35 / dev 60
  397. Task Specialization Fine-Tuning for Contextual Reinforcement Learning imp 50 / dev 75
  398. Reinforcement Learning as (Discrete) Potential Theory imp 35 / dev 70
  399. How smoothing the affinity matrix affects neighborhood preservation in t-SNE imp 25 / dev 60
  400. Pessimistic Meta-Induction and Its Limits: Lessons from Frequentist Statistics and Machine Learning Theory imp 20 / dev 30
  401. Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection imp 20 / dev 20
  402. Physics-Informed and Hybrid Machine Learning in Additive Manufacturing: Application to Fused Filament Fabrication imp 15 / dev 30
  403. Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics imp 50 / dev 60
  404. Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions imp 10 / dev 25
  405. Abra: Scaling Diffusion Image Training imp 65 / dev 65
  406. Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting imp 15 / dev 35
  407. Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements imp 75 / dev 75
  408. MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting imp 10 / dev 40
  409. Tight Bounds for Data-driven Multiple Hyper-parameter Tuning with Structured Loss Function imp 30 / dev 50
  410. Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning imp 45 / dev 60
  411. CORAM: Coherent Orthogonal Rotation for Model Merging imp 55 / dev 70
  412. Pathology Transport: Optimal-Transport Explanations for Clinical Data, and When Their Heatmaps (Fail to) Localize Disease imp 20 / dev 25
  413. Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning imp 45 / dev 65
  414. GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models imp 65 / dev 70
  415. General Semantic Knowledge Infusion for Spatio-Temporal Traffic Forecasting imp 10 / dev 30
  416. Causal Local States: Scalable Simultaneous Causal Network Inference and Forecasting for Dynamical Systems imp 25 / dev 35
  417. Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents imp 55 / dev 65
  418. No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models imp 50 / dev 70
  419. Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries imp 25 / dev 40
  420. OOD Detection for EEG-based Machine Learning in High-Risk Environments imp 25 / dev 35
  421. rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment imp 60 / dev 80
  422. Elimination Geometry imp 15 / dev 35
  423. Picard Proximal Monte Carlo for Parallel Bayesian Imaging with Score-Based Generative Priors imp 20 / dev 40
  424. Conformal Prediction for Molecular Properties under Label Shift imp 25 / dev 40
  425. Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment imp 50 / dev 65
  426. MAGPIE-Net: Predicting short-duration heavy-rainfall events in station neighborhoods from multitemporal FY-4A AGRI observations imp 10 / dev 30
  427. Training-Free Human-in-the-Loop Anomaly Detection via Memory Bank Correction imp 45 / dev 65
  428. Debate Training Reduces Reward Hacking in RLAIF imp 65 / dev 75
  429. Fourth-Moment Geometry of Rademacher Sums imp 5 / dev 10
  430. An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning imp 55 / dev 70
  431. Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs imp 50 / dev 65
  432. MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models imp 30 / dev 45
  433. Efficient Resource Optimization for Split Federated Learning imp 35 / dev 55
  434. Dynamic Compression in Recurrent Networks imp 50 / dev 70
  435. Hybrid ML for Lightweight Pre-Route Delay Estimation in Open-Source IC Design imp 20 / dev 40
  436. Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation imp 55 / dev 70
  437. SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE imp 45 / dev 65
  438. An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models imp 50 / dev 70
  439. Understanding the Surprising Generalization Properties of Tabular Foundation Models imp 45 / dev 65
  440. Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection imp 35 / dev 55
  441. Evaluating and improving crop-yield forecasting methods during extreme drought imp 10 / dev 30
  442. Recirculation imp 55 / dev 75
  443. Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields imp 25 / dev 50
  444. Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents imp 65 / dev 75
  445. Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule imp 10 / dev 35
  446. Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System imp 35 / dev 55
  447. TabNSM: Neural Sparse Mixer for Tabular Regression imp 30 / dev 50
  448. Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization imp 40 / dev 65
  449. The concentration game: Bayesian updating, regret, and information imp 10 / dev 25
  450. Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs imp 50 / dev 70
  451. Information Spreading in Diffusion Models from Effective Field Theory imp 35 / dev 60
  452. Advancing Health Equity through Multi-Level Fairness in Health Informatics imp 25 / dev 35
  453. ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection imp 10 / dev 35
  454. Which CS1 Students Will Fail? Identifying Digital Markers from Learning Analytics in Computer Systems and Architecture Using Weighted Academic Momentum and Interaction Logs imp 10 / dev 25
  455. MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering Model imp 55 / dev 70
  456. Network Denoising Revisited: A Ricci-Flow-Inspired Graph Diffusion Method imp 10 / dev 35
  457. SPSA Hyperparameter Tuning for Variational Quantum Natural Language Inference imp 20 / dev 40
  458. A Constant-Competitive Algorithm for Dynamic Mixture-of-Experts Serving imp 45 / dev 70
  459. WONDER: A Radio World Model-based Negotiation Framework for Multi-Agent UAV Coverage Optimization imp 35 / dev 55
  460. The Price of Thinking: Reasoning Effort as a Model-Specific API Contract imp 60 / dev 75
  461. MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology imp 20 / dev 35
  462. Diagonal Multi-omics Integration of Heterogenous Datasets imp 10 / dev 30
  463. Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations imp 50 / dev 70
  464. FedPref: Federated Preference Learning for Structured Radiology Report Extraction imp 25 / dev 40
  465. Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence imp 20 / dev 35
  466. VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation imp 55 / dev 75
  467. Lambda-Hold Control: Human-Like Movement Emerges from a Minimal Task Reward in Predictive Musculoskeletal Simulation imp 30 / dev 55
  468. Wasted large language models: A life cycle thinking approach imp 40 / dev 35
  469. Dynamic Entanglement-Weighted Pruning for Quantum Federated Unlearning in Supply-Chain Risk Prediction imp 15 / dev 35
  470. Digital Twin-Based Intrusion Detection for Vehicle Powertrain CAN Bus Systems imp 20 / dev 40
  471. Picture the Epsilon: Pursuing Identity-Level Privacy Guarantees for Images imp 40 / dev 60
  472. Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport imp 20 / dev 35
  473. Policy Optimization and Statistical Inference for Online Contextual Matrix Games imp 30 / dev 55
  474. Expressivity In Multimodal Contrastive Learning imp 45 / dev 65
  475. Teach and Grow: An Agent-Centered Architecture for General Robot Learning imp 50 / dev 70
  476. Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal imp 35 / dev 60
  477. Information fusion and machine learning for sensitivity analysis using physics knowledge and experimental data imp 20 / dev 40
  478. Adaptive surrogate modeling for high-dimensional spatio-temporal output imp 20 / dev 40
  479. SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting imp 35 / dev 60
  480. On the Pseudo-Mixing of Kac's Walk imp 5 / dev 10
  481. Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design imp 25 / dev 45
  482. Prism-GRPO: Faster VLA Policy Optimization via Splitting Same-outcome Groups imp 65 / dev 75
  483. Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines imp 10 / dev 35
  484. Online Generalized Sparse Regression: How Does Overparametrization Help? imp 15 / dev 40
  485. When AI Designs AI: Innovation or Imitation? imp 50 / dev 70
  486. SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models imp 50 / dev 70
  487. Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics? imp 20 / dev 35
  488. When to Review: Spaced Repetition for Continual Pre-Training of Language Models imp 55 / dev 70
  489. Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings imp 55 / dev 75
  490. Leveraging existing sparse point annotations for benthic imagery dense segmentation imp 10 / dev 30
  491. Feature Priming in Online Linear Regression: Sparse-Regret Lower Bounds and a Tight Univariate Rate imp 15 / dev 35
  492. Communication Reduction via Semantic-Based Encoding in DMPC Using LSTMs imp 20 / dev 40
  493. MoNe: Modular Neural Memory for Efficient Long Context Inference imp 55 / dev 75
  494. Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision imp 25 / dev 55
  495. Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals imp 55 / dev 75
  496. MemCatalyst: Amplifying Data Auditing on Vision-Language Models via Data Poisoning imp 45 / dev 70
  497. Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See imp 45 / dev 70
  498. Diff-DDoS: Realistic Cyber-Physical Attack Synthesis and Robust Detection for 5G-Enabled CPS Using Tabular Diffusion Models imp 30 / dev 50
  499. Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors imp 10 / dev 30
  500. Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits imp 10 / dev 35
  501. A Residual Learning Approach for Unsteady Aerodynamic Load Prediction imp 15 / dev 25
  502. AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM imp 20 / dev 30
  503. Procedural Content Metageneration via Program Search and Continual Abstraction Discovery imp 45 / dev 60
  504. Towards Zero-Shot Task Transfer with Neurosymbolic World Models imp 50 / dev 55
  505. Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media imp 25 / dev 35
  506. Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry imp 40 / dev 50
  507. Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors imp 20 / dev 25
  508. Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction imp 20 / dev 25
  509. TokEval: A Tokenizer Evaluation Suite imp 55 / dev 75
  510. On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification imp 65 / dev 70
  511. HyPE-GT: where Graph Transformers meet Hyperbolic Positional Encodings imp 35 / dev 45
  512. Diffusion Models for Smarter UAVs: Decision-Making and Modeling imp 35 / dev 45
  513. Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization imp 50 / dev 65
  514. LZ Penalty: An information-theoretic repetition penalty for autoregressive language models imp 55 / dev 70
  515. TabularQGAN: A quantum generative model for tabular data synthesis imp 40 / dev 45
  516. Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures imp 30 / dev 35
  517. Monotone Classification with Relative Approximations imp 25 / dev 30
  518. Continuous Evolution Pool: Taming Recurring Concept Drift in Online Time Series Forecasting imp 50 / dev 60
  519. Scientific Machine Learning of Chaotic Systems Learns Reduced-Order Equations for Neural Populations imp 40 / dev 50
  520. Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification imp 45 / dev 65
  521. Neural Operator-Based Nonlinear Nudging for Chaotic Dynamical Systems imp 40 / dev 50
  522. HeteRo-Select: Informativeness as the Participation Driver in Heterogeneous Federated Learning imp 50 / dev 70
  523. Estimating Parameter Fields in Multi-Physics PDEs from Scarce Measurements imp 40 / dev 50
  524. Asynchronous Message Passing for Addressing Oversquashing in Graph Neural Networks imp 50 / dev 70
  525. One Pipeline, Many Transformers: Pattern-Specific Imputation Specialists for Tabular Missing Data imp 55 / dev 70
  526. A Weak Penalty Neural ODE for Learning Chaotic Dynamics from Noisy Time Series imp 45 / dev 55
  527. Row-Stochastic Matrices Can Provably Outperform Doubly Stochastic Matrices in Decentralized Learning imp 40 / dev 50
  528. Cluster Aggregated GAN (CAG): A Cluster-Based Hybrid Model for Appliance Pattern Generation imp 40 / dev 50
  529. Parametric and Generative Forecasts of EPEX Day-Ahead Energy Market Curves imp 40 / dev 50
  530. How (Not) to Hybridize Neural and Mechanistic Models for Epidemiological Forecasting imp 45 / dev 55
  531. Community Concealment from Graph Neural Networks imp 55 / dev 70
  532. TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts imp 50 / dev 70
  533. Quantifying Memorization and Privacy Risks in Genomic Language Models imp 60 / dev 70
  534. How to make the most of your masked language model for protein engineering imp 50 / dev 65
  535. Likelihood Hacking in Probabilistic Program Synthesis imp 55 / dev 75
  536. Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting imp 65 / dev 80
  537. Protect the Brain When Treating the Heart: Feasibility of 2.5D U-Net for Real-Time Gaseous Microemboli Detection imp 35 / dev 45
  538. The Optimal Sample Complexity of Multiclass and List Learning imp 40 / dev 50
  539. A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning imp 40 / dev 55
  540. Latent Order Bandits imp 40 / dev 55
  541. Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks imp 40 / dev 50
  542. FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers imp 60 / dev 75
  543. ImplicitTerrainV2: Wavelet-Guided Spatially Adaptive Neural Terrain Representation imp 35 / dev 50
  544. Open datasets and machine learning for two-phase heat transfer: a review following a spatial-temporal taxonomy imp 40 / dev 55
  545. Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis imp 50 / dev 65
  546. TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery imp 55 / dev 70
  547. BRo-JEPA: Learning Modular Transformations in Latent Space imp 50 / dev 65
  548. Mos-Gen: A Generative Molecular Framework for Mosquito Insecticide Design imp 45 / dev 60
  549. The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics imp 60 / dev 75
  550. How Transparent is DiffusionGemma? imp 55 / dev 70
  551. Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks imp 45 / dev 55
  552. Low-dimensional topology of deep neural networks imp 40 / dev 50
  553. SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis imp 70 / dev 80
  554. GEqTrain: A Configuration-Driven Framework for Retargeting Equivariant Graph Neural Networks Across 3D Scientific Tasks imp 55 / dev 80
  555. ClockRoPE: Random Fourier Rotations for Temporal Routine Modeling imp 55 / dev 75
  556. A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics imp 40 / dev 55
  557. A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction imp 40 / dev 55
  558. Online Learning of Scale Parameters in Score-Driven Filters imp 40 / dev 55
  559. Geometric and Behavioral Stratification in Transformer Residual Streams imp 55 / dev 70
  560. A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family imp 70 / dev 85
  561. Federated Compositional Muon Optimizer for Matrix-Wise Models imp 50 / dev 75
  562. CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility imp 50 / dev 65
  563. The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference imp 55 / dev 85
  564. CrevasseSeg: A Label-Efficient UAV Crevasse Segmentation Framework imp 35 / dev 50
  565. NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption imp 50 / dev 70
  566. OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations imp 45 / dev 60
  567. A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators imp 40 / dev 55
  568. Deep Learning Based on Generative Adversarial and Convolutional Neural Networks for Financial Time Series Predictions imp 35 / dev 50
  569. The Authenticity Gap in Human Evaluation imp 55 / dev 70
  570. Doubly robust nearest neighbors in factor models imp 40 / dev 55
  571. EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction imp 50 / dev 70
  572. Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses imp 40 / dev 55
  573. Spikformer V2: Join the High Accuracy Club on ImageNet with an SNN Ticket imp 50 / dev 75
  574. Predicting Male Domestic Violence Using Explainable Ensemble Learning and Exploratory Data Analysis imp 35 / dev 40
  575. On Stability in Optimistic Bilevel Optimization imp 40 / dev 55
  576. Efficient Dynamic Shielding for Parametric Safety Specifications imp 65 / dev 80
  577. On detection probabilities of link invariants imp 5 / dev 0
  578. SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA imp 65 / dev 80
  579. Inverse Problems for Partial Differential Equations with Jump Discontinuities in Coefficients via Two-Stage Physics-Informed Deep Learning and Statistical Mixture Models imp 45 / dev 60
  580. A multi-view contrastive learning framework for spatial embeddings in risk modelling imp 45 / dev 65
  581. SparsePixels: Efficient Convolution for Sparse Data on FPGAs imp 50 / dev 80
  582. Large Language Models: A Mathematical Formulation imp 65 / dev 85
  583. Fermi-Dirac thermal measurements: A framework for quantum hypothesis testing and semidefinite optimization imp 30 / dev 25
  584. Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? imp 65 / dev 80
  585. From Diffusion to Flow: Efficient Motion Generation in MotionGPT3 imp 50 / dev 75
  586. Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models imp 60 / dev 75
  587. Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries imp 55 / dev 70
  588. SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation imp 45 / dev 65
  589. FairNVT: Fair Classification via Noise Injection in Vision Transformers imp 55 / dev 75
  590. Convergent Evolution: How Different Language Models Learn Similar Number Representations imp 55 / dev 75
  591. Nonlinear GENERIC-Embedded Neural Networks (N-GENNs): Learning GENERIC dynamics with non-quadratic dissipation potentials imp 45 / dev 65
  592. Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum imp 60 / dev 80
  593. Memory by Design: Probabilistic Sequence Layers imp 55 / dev 75
  594. SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails imp 40 / dev 65
  595. Spectrally Safe Neural Operator Warm-Starts for Large-Scale Newton Solvers imp 50 / dev 70
  596. Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment imp 65 / dev 85
  597. CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification imp 40 / dev 55
  598. From Adoption to Deployment: A Qualitative Study on AI Integration in Software Development Practice imp 70 / dev 85
  599. Constitutional Midtraining: Content Presence Drives Alignment Gains imp 70 / dev 85
  600. Non-KKT Accumulation in Entropic Mirror Descent imp 35 / dev 50
  601. Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling imp 35 / dev 25
  602. Efficient Hessian-Free Methods for Multi-Objective Bilevel Optimization with Nonconvex Lower Level imp 25 / dev 20
  603. Belayer: Efficient Fault Tolerance for LLM Agentic RL Training imp 65 / dev 65
  604. RecurrentGPT: Expressive Depth through Recurrent Modulation in Transformers imp 45 / dev 40
  605. The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT imp 55 / dev 50
  606. Demo: tfdrift - A Severity Taxonomy and Risk Classification Framework for Infrastructure Drift Detection imp 50 / dev 60
  607. Reproducibility is Not Enough: Artifact Verifiability in Decentralized-Build Package Ecosystems imp 55 / dev 60
  608. Engine-Transfer-Bench: An Evidence-Based Benchmark for Document Compilation Engine Selection imp 30 / dev 40
  609. Building real-time digital twin instances with Function+Data Flow: user evaluation and extension for iterative pipelines imp 40 / dev 50
  610. SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation imp 60 / dev 70
  611. AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin imp 60 / dev 70
  612. OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language imp 45 / dev 60
  613. Code Health in LLM-Based Test Generation: Effectiveness and Token Efficiency imp 55 / dev 65
  614. Contract-Aware Rescue of a Drifted Isabelle Development: The Double-Tank Case Study imp 50 / dev 55
  615. FPGA Lifecycle Management for RISC-V Systems imp 35 / dev 40
  616. MicroPython and CircuitPython: Pythons Quiet Takeover of IoT and Robotics imp 45 / dev 50
  617. When Do Microservices Save Energy? Evidence from Environmental Simulation Workflows imp 40 / dev 45
  618. SiNMULI: Novel Signed Network Approach for Malicious URL Identification imp 40 / dev 40
  619. Toward Inclusive AI-Driven Development: Exploring Gender Differences in Code Generation Tool Interactions imp 50 / dev 50
  620. A Configuration-First Framework for Reproducible, Low-Code Machine Learning: a Localization Use Case imp 45 / dev 55
  621. When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling imp 75 / dev 75
  622. From Quality Properties to Practice: A Guideline and Workflow for Explainability Requirements imp 55 / dev 55
  623. Do Privacy Policies Match with the Logs? An Empirical Study of Privacy Disclosure in Android Application Logs imp 50 / dev 55
  624. MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox imp 65 / dev 80
  625. From Documentation to Zero-day Vulnerabilities: LLM-Driven Fuzzing of JavaScript Engines in PDF Readers imp 65 / dev 70
  626. PowderLine: a programmatic powder diffraction analysis application imp 25 / dev 30
  627. Supply chain attack on arrayref imp 70 / dev 65
  628. What Zig felt like, coming from Rust imp 35 / dev 60
  629. How a joke domain purchase turned into geopolitical warfare imp 30 / dev 10
  630. HTML Can Do That imp 40 / dev 70
  631. Reclaim the terminal imp 35 / dev 60
  632. Open-sourcing OpenPubkey SSH (OPKSSH): integrating single sign-on with SSH imp 55 / dev 70
  633. Plain Text Accounting is Pretty Cool imp 30 / dev 35
  634. AliExpress keeps multipoint Bluetooth headphones active with WebAudio fingerprinting imp 55 / dev 50
  635. Why compiling Rust to WebAssembly is slow imp 45 / dev 70
  636. If this is true, the hyperscalers are toast imp 50 / dev 50
  637. Everyone Says Assembly Is Untyped—Everyone Is Wrong imp 30 / dev 50
  638. Sing-song: a speakable encoding for long numbers and keys imp 35 / dev 40
  639. Opus is a minimal, statically-scoped Lisp dialect based on the semantics of f-expressions (the Kernel language) imp 25 / dev 40
  640. Why I still hand write my commit messages imp 30 / dev 40
  641. X.Org Server 26.1 RC1 Prepares For First Feature Release In Five Years imp 35 / dev 40
  642. Reverse-engineering Find My People to stalk ̶m̶y̶ ̶e̶x̶ a friend, cause I can imp 40 / dev 50
  643. Ploopy A+ (external trackball) imp 15 / dev 15
  644. Understanding the limitations of Pubsub systems imp 50 / dev 70
  645. SQLite for Everything imp 55 / dev 70
  646. Going freestanding imp 35 / dev 60
  647. A Personal Computer For Children Of All Cultures imp 25 / dev 30
  648. Introducing Microlighter imp 25 / dev 30
  649. Unlocking the Future: A Deep Dive into Agentic AI and Autonomous AI Agents imp 60 / dev 70
  650. Top Platforms for AI Guardrails Implementation to Secure Your AI Apps in 2026 imp 60 / dev 65
  651. Tokenization, Training Stages, and Why Bigger Models Work (LLM Internals, Part Three) imp 55 / dev 60
  652. Antes de ampliar data centers, a Dropbox tenta extrair mais da frota atual imp 45 / dev 55
  653. Enterprise AI Security Platforms: Top Choices for 2026 imp 60 / dev 60
  654. Portable Text Classification Backends: 50 JSON Tagging Trials Across Europe and US Apps imp 60 / dev 75
  655. Day 2: Comprehensions & My First Generator-Based File Reader | Python | Day 2 | Week 1 | 10 Week Challenge imp 15 / dev 30
  656. Adaptive RAG: Designing Retrieval Pipelines That Choose the Right Strategy at Runtime imp 70 / dev 80
  657. FeedBoss Review 2026: Features, Pricing & Top Alternatives imp 20 / dev 20
  658. Top 5 Open Source LLM Gateways to Control Your LLM Traffic in 2026 imp 70 / dev 80
  659. I Challenge You: Give Me a Project. I'll Build It in 48 Hours. 🚀 imp 15 / dev 25
  660. Cronloop AI Guide: How to Use It, Best Prompts & Use Cases (2026) imp 55 / dev 60
  661. How we cut repo-wide symbol indexing for LLM agents from 30s to 98ms imp 75 / dev 85
  662. Implementing Healthtech LLM Classification in Node.js: Structured JSON Batch Tagging imp 50 / dev 70
  663. Chapter 3 Core System Components and Internal Implementation imp 45 / dev 65
  664. Building TonuAI: A Fast AI System for E-Commerce imp 50 / dev 65
  665. Batch-Tagging a Messy Game Catalog CSV Without Locking Into One LLM Classification API imp 55 / dev 75
  666. Every release makes the harness harder to fool: LLMKube 0.9.19 imp 70 / dev 80
  667. Why LLMs Hallucinate and How to Reduce Hallucinations imp 70 / dev 75
  668. Async LLM API Jobs for Bulk Invoice Extraction (Beyond Realtime Cost) imp 60 / dev 75
  669. Claude Code with any model: three ways to route it (incl. the 2-minute one) imp 65 / dev 75
  670. The Dashboard Liar's Club: Why Your AI Observability Tools Show Fake Data (And Why That Matters) imp 65 / dev 75
  671. How Unfiltered Neural Models and Aimour AI Are Re-Engineering Digital Intimacy in 2026 imp 20 / dev 20
  672. Discussion Hub for new Claude incident: Elevated errors on Google connectors on Aug 20, 2026 imp 60 / dev 50
  673. This is letting Claude handle a good amount of money for a month... imp 30 / dev 20
  674. How big of a difference do you think it’s going to make on token consumption? imp 35 / dev 40
  675. The Claude language calibration issue on GitHub got an official response from Anthropic. Guess who wrote it. imp 40 / dev 45
  676. PSA: a malicious published Claude artifact is ranking on Google for Claude Code install queries — it installed a macOS infostealer on my Mac imp 80 / dev 70
  677. Claude says I used 54.9 BILLION tokens. imp 20 / dev 15
  678. I Am Morally Opposed to Updating My CLAUDE.md imp 40 / dev 50
  679. Claude is a thinking partner. Opus 5 is not Claude. imp 50 / dev 60
  680. I adapted The Elements of Style to make Claude Code write in plain English imp 55 / dev 60
  681. Hot take: Claude code should ask more questions before touching your code. imp 50 / dev 60
  682. How is the dude who said he will use fable on the quest to get a wife doing? imp 5 / dev 0
  683. Claude Code Weekly Limits imp 35 / dev 40
  684. I'm so careful to not share secret keys with claude imp 50 / dev 55
  685. Claude Sonnet 5 shifts behavior when it recognizes the user as an AI safety researcher imp 60 / dev 65
  686. that feeling when using other models and not having my model get gimped just because i used the word "blood" or "drugs" in my prompt imp 45 / dev 50
  687. Claude: it’s a desktop, ya dingus! imp 20 / dev 25
  688. Show us what you've created with Claude! imp 15 / dev 20
  689. Give Back Claude’s ‘Thought Process’ imp 45 / dev 50
  690. claude started hallucinating imp 30 / dev 35
  691. I read through the new concise output style system prompt, and found a new output style as well. imp 55 / dev 60
  692. Antrophic Employee said there is "make a lot of money" button imp 20 / dev 15
  693. Suddenly getting a bunch of Fable safety guarding on super normal stuff like PR creation imp 50 / dev 55
  694. Knowing exactly what you want and having no idea how to ask for it imp 50 / dev 60
  695. What can I do with Claude Pro before my usage resets? imp 25 / dev 30
  696. exactly the kind of problem AI was made for imp 30 / dev 35
  697. anyone got a problem of storage space getting full and not knowing what to delete imp 20 / dev 25
  698. Claude Code vs. OpenAI Codex for coding ($100 budget) — which offers better value, or is there a better alternative? imp 50 / dev 70
  699. How should a complete beginner validate and build a social app with AI coding tools? imp 50 / dev 65
  700. Programmer Help with Program needed imp 20 / dev 30
  701. Why Reddit is the best social network for developers - and maybe for other people too imp 50 / dev 70
  702. Anyone NOT on full auto when coding with local LLMs? imp 30 / dev 60
  703. Claude Code and Codex on one keyboard — every session gets a lane on the RGB F-row, with a summon key per agent imp 25 / dev 55
  704. How would you structure an AI-assisted React Native rewrite workflow? imp 40 / dev 75
  705. We clicked 48 AI-generated web apps in a real browser — the pricier model failed more than the cheap one imp 65 / dev 70
  706. Is OpenAI using Chinese LLM models for their service, not theirs? imp 15 / dev 20
  707. I just built a mini Kimi-K3 from Scratch under 250$. Already beats GPT-2 (124M)! imp 35 / dev 75
  708. Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant imp 30 / dev 70
  709. Ling-3.0 released all 6 base checkpoints: 2 sizes × 3 stages imp 50 / dev 75
  710. The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches imp 35 / dev 75
  711. QwenMix-3.7: Kept seeing posts about Qwen3.8 and 3.6 sharing the same structure.. so I had Qwen3.8 combine them. imp 25 / dev 70
  712. Tencent begins testing its new flagship model Hunyuan Hy4 imp 65 / dev 75
  713. Qwen3.8-27b has the highest level of "agency" I've ever seen in a local model imp 50 / dev 80
  714. Aurora-80K releases! A modern tiny language model. imp 40 / dev 70
  715. Getting better at coding doesn't make a model better at everything else imp 45 / dev 65
  716. AQuA's "self-improvement" updates research state, not the agent LM. What should a local port freeze? imp 55 / dev 75
  717. New benchmark just dropped! imp 35 / dev 70
  718. [MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model! imp 40 / dev 75
  719. Theres surely SOMEONE out there whose job is just pumping out low-poly oneshot ThreeJS assets.. imp 15 / dev 20
  720. Qwen3.8-27B took a serious hit to *knowledge* vs 3.6 imp 50 / dev 75
  721. Qwen3.8 27b just exceeded my expectations on svg generation :D imp 35 / dev 70
  722. TinySearch v0.6.1 - still a lightweight web research tool for local LLMs, now with bring-your-own-browser support imp 45 / dev 80
  723. [Draft - Open PR] AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp imp 50 / dev 85
  724. Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs imp 45 / dev 80
  725. GLM 5.3 SlopCodeBench Results imp 55 / dev 80
  726. Spider-man: Brand New Day, does Peter self host his AI? (Spoilers) imp 25 / dev 50
  727. Claude sonnet 4.6 was really good at estimating the future qwen 3.8 27b performance imp 35 / dev 70
  728. G9v3-39A5B on artificialanalysis looks good. Has anyone tested it? imp 20 / dev 60
  729. AirLLM - Recent Updates - with Qwen3.8-27B, Kimi-K3 too imp 45 / dev 75
  730. Cleanest way you've found to A/B two models in the same agent? imp 55 / dev 75
  731. Unpopular opinion: AI agents don't always need a Vector DB for project memory imp 55 / dev 80
  732. An AI agent isn’t production-ready until a human can take over halfway through a run imp 60 / dev 80
  733. Running one voice agent across multiple countries is way messier than running one per market. How are people handling it? imp 55 / dev 75
  734. What Does Reddit Think: What’s one boring AI capability that would completely change your work? imp 40 / dev 60
  735. What matters more for an entrepreneur: knowing AI tools or knowing how to implement AI into a real business? imp 30 / dev 50
  736. What modes does your agent have besides Plan Mode? imp 50 / dev 75
  737. I built a governance layer for CrewAI (pip install crewai-governance) imp 60 / dev 85
  738. We armed auto-merge on 108 agent-written pull requests. One merged. imp 65 / dev 85
  739. Best open source calendar. imp 10 / dev 40
  740. How and How Often Are you Re-evaluating agent value? imp 50 / dev 70
  741. The agent failures that get you aren't crashes. They're clean runs that did the wrong thing. imp 70 / dev 85
  742. Seeing a lot of people post about on maintaining context across various AI providers and chats, here's a tool to help you. imp 45 / dev 75
  743. What is one business problem you think AI agents are actually good at solving? imp 50 / dev 65
  744. Have you tried any open source harness similar to claudes's managed agents but costs less? imp 70 / dev 85
  745. How we use an AI desktop agent to lock in brand consistency across multi-asset campaigns imp 55 / dev 75
  746. Lightweight memory for agents imp 55 / dev 80
  747. Looking for an Agentic AI Job/Interview Opportunity as a Fresher imp 15 / dev 40
  748. Casi 22 días, un solo objetivo y un repositorio de 1,8 millones de líneas: ¿estamos midiendo mal la autonomía de los agentes? imp 70 / dev 85
  749. Are we giving AI agents too much autonomy too early? imp 60 / dev 80
  750. I built an open-source interoperability layer for AI agents imp 70 / dev 85
  751. Make Web-Sites actionable for arbitrary AI Agents imp 60 / dev 80
  752. Gave my coding agents SSH access to real servers without putting keys in their environment - here's the trust model imp 70 / dev 85
  753. New age insults imp 5 / dev 5
  754. A Deluge of A.I. Computing Power Is About to Come Online, Fueling Major Leaps | The number of A.I. chips that provide the computing power to advance the fast-evolving technology is doubling every nine months. imp 75 / dev 75
  755. Did GPT 5.6 Sol get secretly upgraded? imp 20 / dev 40
  756. If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective) imp 35 / dev 45
  757. POV: you're born as an AI imp 5 / dev 5
  758. Anyone else having this issue? imp 20 / dev 40
  759. I made a notetaker that runs on your ChatGPT subscription imp 40 / dev 70
  760. I asked AI to take a random photo with an iPhone 6 flash imp 15 / dev 50
  761. Built a stateful AI D&D & World Simulation engine with 1-Click Campaign Sharing — Test our House of the Dragon campaign imp 45 / dev 70
  762. Tibo, where are you when I need a reset imp 20 / dev 35
  763. Voice mode got weird imp 15 / dev 30
  764. Vent: Prompts for building a Bluetooth Sink for Audio keep getting flagged imp 25 / dev 50
  765. Unable to Log in, Is it just me? imp 10 / dev 20
  766. ChatGPT is down imp 10 / dev 20
  767. OpenAI halts testing, slows development after rogue model hacked Hugging Face imp 75 / dev 80
  768. Sol 5.6 model suddenly become very dumb imp 30 / dev 50
  769. Researchers created "mind viruses" that spread between AI agents by convincing one agent to adopt an idea then transmit it onwards to other agents. imp 75 / dev 85
  770. chatgpt really struggling... imp 25 / dev 40
  771. I’m surprised everyone is okay with the Tibo reset era imp 20 / dev 35
  772. Does OpenAI makes you scream? imp 5 / dev 5
  773. What a sales pitch! imp 5 / dev 5
  774. Is ChatGPT classic a dead end? imp 30 / dev 50
  775. This is my new record imp 20 / dev 50
  776. smolmachines / smolvm as a sandbox for untrusted Python & JavaScript imp 65 / dev 85
  777. Quoting Jeremy Morrell imp 50 / dev 75
  778. Conceptual integrity and counting lines of code imp 50 / dev 70
  779. [AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law imp 75 / dev 85
  780. [AINews] Memory prices up 500% in 12 months imp 70 / dev 75
  781. Grok 4.6 imp 70 / dev 80
  782. Lifelong imp 15 / dev 35
  783. MeetStream AI imp 50 / dev 75
  784. Prized imp 40 / dev 70
  785. The New Calendly imp 30 / dev 55
  786. Hermai Brand API imp 25 / dev 60
  787. Glasp for Firefox imp 25 / dev 60
  788. Aloud imp 45 / dev 75
  789. ProtoNote imp 30 / dev 65
  790. Cloudways Managed AI Agents imp 50 / dev 75
  791. NobodyWho imp 40 / dev 75
  792. Peach Co-Pilot imp 30 / dev 65
  793. MiniMax Design imp 50 / dev 75
  794. Shape imp 55 / dev 80
  795. Revy imp 15 / dev 40
  796. Fairphone Gen 6+ imp 15 / dev 40
  797. KiHub imp 15 / dev 50
  798. Basedash Public Sharing imp 30 / dev 65
  799. Loopcase imp 20 / dev 55
  800. Vois 2.0 imp 40 / dev 75
  801. Claude Watermark Remover imp 15 / dev 30
  802. Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier imp 0 / dev 70
  803. Sokoban via Grok App Builder imp 10 / dev 25
  804. Oh My OpenCode Slim imp 35 / dev 65
  805. Evaluating DeepSeek V4 Pro 0813 on Hack the Box Challenges imp 25 / dev 45
  806. Show HN: LLM-as-a-Verifier Plugin for DeepSeek Harness imp 50 / dev 75
  807. Run GLM-OCR, DeepSeek-OCR-2, Dots.mocr with an OpenAI Compatible API imp 35 / dev 60
  808. Show HN: Open Bot – an open-source Grok Bot that works with any agent harness imp 65 / dev 75
  809. "the mandate" a short fiction by Grok 4.6 on technological singularity imp 5 / dev 5
  810. DeepSeek Hikes AI Prices More Than 4x, What It Means for NVDA, MSFT, AMZN and BABA - TipRanks imp 55 / dev 25
  811. US Lead in the AI Race With China Is Rapidly Narrowing - Bloomberg.com imp 40 / dev 20
  812. 😻 AI Tool Roundup: Qwen, Cursor, DeepSeek Explained Live - The Neuron imp 25 / dev 40
  813. 2Q26 Datacenter Supply/Demand: DeepSeek to Claude - CreditSights imp 35 / dev 30
  814. How to Set Up DeepSeek V4 Pro: 12 Steps, 90 Min [2026] - tech-insider.org imp 20 / dev 55
  815. Chinese AI firms forced to optimise software as local chips trail Nvidia’s - South China Morning Post imp 50 / dev 65
  816. GPT-5.6 vs DeepSeek V4 Pro 0813: 714x Cheaper Input [2026] - tech-insider.org imp 30 / dev 40
  817. How I Reached #7 in a Hugging Face AI Competition for Under $80 - HackerNoon imp 20 / dev 40
  818. Pushing or Pacing or Pulling Back From the AI Frontier? - spyglass.org imp 30 / dev 15
  819. Why China Is Approving NVIDIA H200 AI Chip Imports? - AI Magazine imp 25 / dev 25
  820. The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard - TheSequence | Jesus Rodriguez imp 45 / dev 70
  821. Best AI Chatbots 2026: ChatGPT vs Claude vs Gemini - tech-insider.org imp 20 / dev 35
  822. Snowflake lets Cortex AI gateway choose models itself - Techzine Global imp 55 / dev 75
  823. OpenAI, Claude, Gemini or DeepSeek – Which AI Is Better at Discovering Drug Candidates? - KTBS 3 imp 35 / dev 45
  824. DeepSeek V4 Pro vs Qwen 3.8 Max: Pricing and Open-Weight Changes Shift the Comparison - Memeburn imp 30 / dev 45
  825. Tussle between open and closed AI reshaping global order - China Daily imp 40 / dev 20
  826. Does One Unique Skill Make DeepSeek V4 Pro Outperform Fable 5? The Viral AI Plugin Myth Debunked - 36Kr imp 35 / dev 60
  827. Top 10: AI Leaders in APAC - AI Magazine imp 20 / dev 15
  828. Snowflake Adds Dynamic Model Routing to Cut Enterprise AI Costs - analyticsindiamag.com imp 55 / dev 70
  829. Moonshot AI’s IPO needs a new story after Kimi K3 - KrASIA imp 25 / dev 15
  830. Wang Xingxing Enters the "Odyssey" Period, Sincerely Calls for Industry Partner Liang Wenfeng - 36Kr imp 15 / dev 10
  831. Zhipu's Comeback Code: Betting on Coding Delivers a HK$500 Billion Valuation and $1 Billion ARR - finance.biggo.com imp 25 / dev 20
  832. Grok Can Earn You Money: What Musk's Claim Actually Means - BASENOR - Tesla Accessories imp 20 / dev 15
  833. Grok Bot Runs on a Remote Computer — Even When Yours Is Off - BASENOR - Tesla Accessories imp 45 / dev 65
  834. Grok predicts Notre Dame to have a massive improvement in a crucial area ahead of the 2026 season - A to Z Sports imp 2 / dev 5
  835. xAI Launches Grok Build: An Agentic CLI That Runs Your Computer - BASENOR - Tesla Accessories imp 75 / dev 85
  836. How Cybercriminals Are Weaponizing Frontier AI Models Like Grok - Forbes imp 60 / dev 45
  837. From energy to AI: Five major US business deals of 2026 - The American Bazaar imp 20 / dev 10
  838. Silicon Valley CEO: The Birth of Grok Bot Is As Shocking As the Claude Code Moment - 36Kr imp 50 / dev 70
  839. E&E News: Memphis weighs data center moratorium in city Musk’s xAI calls ‘home’ - POLITICO Pro imp 15 / dev 5
  840. Grok Adds App Sharing and Access Controls for Builders - BASENOR - Tesla Accessories imp 35 / dev 50
  841. James May criticizes Tesla’s Grok AI voice assistant as ‘insincere and phony’ - eciks.org imp 10 / dev 5
  842. Grok Can Now Control Smart Home Devices Remotely - BASENOR - Tesla Accessories imp 30 / dev 45
  843. X offers $175K for Grok-generated versions of The Odyssey - Social Media Today imp 10 / dev 5
  844. Elon Musk Teases Grok Voice — What We Know So Far - BASENOR - Tesla Accessories imp 15 / dev 10
  845. Elon Musk Grok AI Predicts Gold Price Will Explode by End of 2026 - 99Bitcoins imp 10 / dev 20
  846. Grok 4.6 Hits #1 for Healthcare Questions - BASENOR - Tesla Accessories imp 25 / dev 40
  847. 【レポート】田町.ai #2 LT大会を開催しました imp 10 / dev 25
  848. LLM導入で失敗する原因はモデルだけではない:開発前に確認したい4つの層 imp 55 / dev 75
  849. AI AgentがTool選択で失敗する理由:MCP設計で見直したい5つのポイント imp 70 / dev 85
  850. AI AgentによるX投稿文の生成 imp 25 / dev 50
  851. GPUメモリの壁を進化戦略で越える — Agentic ESOptが示す長期エージェント学習の現実解 imp 65 / dev 80
  852. RAGと何が違う?AIエージェントに同一性と忘却を与える複合記憶アーキテクチャ imp 70 / dev 85
  853. # AIが「できたか分からない」と言ったとき、もう一度やらせてはいけない## ―― 外部API・二重実行・Commit-Unknownか imp 75 / dev 80
  854. GRPOからDAPOへ:RLVR時代のLLM強化学習を数式とメカニズムで理解する imp 50 / dev 75
  855. ローカルLLM本番運用フルスタック:vLLM・SREの最小構成 imp 65 / dev 85
  856. 止まらないエージェントを実行前に見つける静的解析ツールIAL-Scan imp 70 / dev 85
  857. Claude Codeの回答を"行動指向"にするOutput Styleは効果があるのか検証してみた imp 45 / dev 65
  858. Agent Skillsは「なぜ」効いているのか — 検索精度はスケールで崩壊する imp 65 / dev 80
  859. AIエージェントに「記憶」を持たせるべきか。文脈が累積するかで決まる imp 60 / dev 75
  860. NeoBrowser(実Chrome操作のMCPサーバー)を動かしたら目玉機能が非対話環境で死んだ imp 45 / dev 80
  861. AIエージェントの評価は「1回合格」では足りない ― 敵対テスト×トレース×統計リプレイで測ると60%失敗 imp 70 / dev 85
  862. 「そこを疑え」は人間にしか出せない、と書こうとして3つ穴が空いた imp 55 / dev 70
  863. 【最新AI/開発ツール】Unsloth Dynamic 3.0 GGUF解説と注目ニュース imp 50 / dev 80
  864. 多役プロンプトの後半ロール希薄化は、守るものの記述量で緩和できるか検証した imp 35 / dev 60
  865. 相手の AI に「聞き方」を送ることにした (LLM が両端にいるのに、真ん中で人間が文章を運んでいる問題) imp 55 / dev 75
  866. YANS2026 参加報告 imp 10 / dev 25
  867. 文書構造解析を検索の前処理にする:壊さないパース imp 55 / dev 80
  868. LLM・生成モデルの推論高速化技術の全体像(多分) imp 65 / dev 85
  869. はんなりPython#12でLTをしてきました imp 0 / dev 10
  870. パーコレーションの相転移を実測する。占有率1.5ポイントの差で貫通確率が0%から100%に跳ぶ imp 0 / dev 5
  871. ニューラルネットワークとは?脳を模した学習モデル imp 15 / dev 40
  872. クラシック・データサイエンスの現在地 #1 Bootstrap — 統計学が「計算」を手に入れたとき imp 10 / dev 30
  873. 第6回 関東Kaggler会 参加レポート imp 10 / dev 25
  874. RL ポストトレーニングが静かに壊れるとき: training と inference の数値ミスマッチ imp 60 / dev 80
  875. Seed-VCアーキテクチャ徹底解説 — 声を「誰が・何を・どう」に分解する4段構成 imp 30 / dev 70
  876. 【今日から俺もFDE #2】ChatGPTは「うちの会社」を知らない — 社内ChatBotを最短で作る方法(後編) imp 40 / dev 65
  877. 議事録も実験結果も全部mdに ── 松尾研究所の半分以上のプロジェクトで動く「LLM Wiki」を勉強会で覗いてきた imp 65 / dev 75
  878. 【さくらのAI】🔰さくらのAIを使って、はじめてのLLMゲームを作ろう【請求0円/3000回】 imp 20 / dev 45
  879. プロンプトインジェクションの「その後」を設計する — エージェントフレームワークの信頼境界 imp 60 / dev 80
  880. 因果推論 Day 9/全30回 感度分析とE-value、隠れた交絡にどこまで耐えるか imp 15 / dev 40
  881. 【技術解説】【完全ガイド】Pythonで松井証券の自動売買を実現する方法と実践コード imp 15 / dev 40
  882. プロンプトインジェクション検知の精度は言語によって変わるのか?6言語・6000文で検証してみた imp 50 / dev 75
  883. 医用画像(DICOM/NIfTI)のPNG変換と機械学習の基礎的アプローチ【備忘録】 imp 20 / dev 55
  884. 月の石 imp 0 / dev 0
  885. 崖・蜃気楼・台帳——LLM創発研究の現在地を、当事者のひとりが検分する imp 50 / dev 60
  886. LLMにLLMを評価させてみた imp 40 / dev 70
  887. 源内の中で動いているのはAnthropicとAWSのモデルで、公表されている効果は1,200人分の数字だった imp 45 / dev 50
  888. 見られることと、理解されること imp 0 / dev 0
  889. マルチAIを役割分担で使ってる私が、Anthropicの新研究を見て思ったこと imp 45 / dev 70
  890. Claude Opusは無能な働き者である imp 15 / dev 30
  891. 【続編】AIっぽいところ(3か所)、AIしくじってるなというところ(3か所)『、、、僕がやったのは、パスワードを入力したことくらい。、、、#松浦勝人』 #松浦勝人 imp 10 / dev 20
  892. 雰囲気で終わらせないローカルLLM用語解説 imp 35 / dev 75
  893. AIよもやま話 #004|AIに自分を理解させるということ imp 35 / dev 55
  894. シンガポール・コンセンサス2026 / 予防から社会的回復力へ / エージェント十原則と責任分界 雑感 imp 55 / dev 50
  895. Claude Codeとローカルqwenの分業ライン imp 40 / dev 75
  896. AIを倒錯紳士にしたら、モデルごとの「性癖」が見え始めた話 imp 15 / dev 30
  897. 【生成AIニュース+】『MAI-Image-2.6』『Meshy 7』『ComfyUI-MiniMaxH3-Parallel』『Marketing OS by Arcads』『Tripo P2.0 Preview』『Raon-OpenTTS-1B』『MinimaxH3_Characters』『LTX2.5_actions』『ReelBids LTX-2.5 Camera LoRA』『ComfyUI-Raon-OpenTTS』『MiniMax-H3 Turbo』他多数 imp 40 / dev 70
  898. 13万円のGPUを買う直前でやめた話 imp 20 / dev 50
  899. LLMにも効く「アンカリング効果」──AnchorBenchが教える運用上の注意点 imp 50 / dev 70
  900. 【GPT】人間の皮膚を完全再現。Realized Hybrid Systemに⑤これだけプロンプトを追加‼️パーソナライズ設定×GPT×プロンプト imp 15 / dev 40
  901. 【とりあえず勝ち越し】MT5 LLM自動売買Bot6機 並走トレード録 8/19 imp 5 / dev 15
  902. 生成AIはなぜ「知らない」と言えず嘘をつくのか──音喜多駿の肩書き誤認から見る、ハルシネーション対策の現在地 imp 65 / dev 75
  903. ChatGPTのメモリだけに頼らない。だから「引き継げる相棒」を作った。 imp 45 / dev 65
  904. LLMの臨床エラー検出評価に新提案:F1値の落とし穴とペア評価の重要性 imp 55 / dev 75
  905. LLM#1 temperature 0.7 の意味を、私は説明できなかった imp 40 / dev 85
  906. もしかしたら特定世代のオープンウェイトモデルが貴重になるかも… imp 30 / dev 55
  907. 【第8話】AIに仕事を取らせるな、「整理」をさせよ——一人社長のデスクワークを自動化するタスク抽出ベンチマークと業務フロー設計 imp 55 / dev 60
  908. Fluter 3.47正式リリース。UIライブラリが分離され独立してアップデート可能、デフォルトでWebAssemblyを生成する方向に、など新機能 imp 5 / dev 10
  909. Docker社、コンテナ向けの高性能な新ハイパーバイザ「Docker VMM」パブリックベータ公開 imp 25 / dev 40
  910. Pythonライクな新言語「Mojo」がオープンソースで公開、コンパイラやツールチェーンなど。Windows版の開発も表明 imp 70 / dev 75
  911. メルカリにおけるTiDB改善の取り組み:インデックス非互換への対応 imp 30 / dev 65
  912. メルカリにおけるTiDB改善の取り組み:リソース制御・リージョンサイズ・プランキャッシュ編 imp 30 / dev 65
  913. 2026年9月の技術系イベント予定 imp 5 / dev 10
  914. AIを入れても、なぜ業務量は思ったほど減らないのか imp 50 / dev 55