AI News Digest 2026-09-02
台本で使った記事
特集
開発者コーナー
中堅コーナー
ハーネスコーナー
速報コーナー
参考記事一覧
参考記事一覧を表示(1322件)
- Introducing Claude Fable 5.1 and Claude Mythos 5.1r/ClaudeAI / imp 75 / dev 60
- Department of War Launches Starshield AI's Grok for Government on GenAI.mil - U.S. Department of War (.gov)Google News Grok/xAI / imp 45 / dev 20
- Elon Musk calls Grok ‘boring’ as AI roasts Billie Eilish over capitalism comments - FirstpostGoogle News Grok/xAI / imp 5 / dev 0
- xAI trained Grok on child sex abuse material, a new class action lawsuit alleges - claimdepot.comGoogle News Grok/xAI / imp 60 / dev 15
- Tesla’s Grok can now control more than 100 vehicle functions hands-free [List] - driveteslacanada.caGoogle News Grok/xAI / imp 35 / dev 30
- MoonPay Turns Grok Into a Chat Interface for Buying and Lending Crypto - Startup FortuneGoogle News Grok/xAI / imp 30 / dev 35
- Instagram admits users often can't tell AI profiles from real peopleThe Decoder / imp 50 / dev 25
- ChatGPT now faces stricter EU oversight as a very large search engineThe Decoder / imp 55 / dev 20
- Google Pics is like Canva, but with even more AIThe Verge (AI) / imp 45 / dev 50
- Healthcare organizations can now connect EHR and additional industry data to ChatGPTOpenAI News / imp 40 / dev 45
- Apple accuses OpenAI of destroying evidenceThe Verge (AI) / imp 50 / dev 10
- Sony and Warner just sued Anthropic for the exact same piracy Anthropic already admitted to and paid $1.5B forr/artificial / imp 55 / dev 10
- DeepSeek Open-Sources First V4 Vision Model: Benchmark Claims Need Independent Proof - techtimes.comGoogle News DeepSeek / imp 65 / dev 70
- Aurora Mobile’s Modellix Releases Beta Plugin for DeepSeek Harness, Adding Free LLM Models to the Fast-Growing Open-Source Coding Agent - The Manila TimesGoogle News DeepSeek / imp 65 / dev 75
- China AI Model Monitoring: Accelerated Upgrades, Surge in Usage, and Cloud Entering a New Upward Cycle - MoomooGoogle News DeepSeek / imp 40 / dev 35
- Keenable SELECT: an agent that searches the web in SQLr/LocalLLaMA / imp 55 / dev 75
- It must be some kind of psy-op by OpenAI to claim that Sol is anywhere near as good as Fabler/ClaudeAI / imp 15 / dev 50
- Play Store blocks AuroraStore, hurting GrapheneOS usersHacker News / imp 35 / dev 40
- The creator of Jujutsu has joined ERSCHacker News / imp 45 / dev 65
- Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for developmentHacker News / imp 40 / dev 55
- AnkiDroid: Google Play no longer allowing Open Collective donation linkHacker News / imp 35 / dev 30
- Ask HN: Who is hiring? (September 2026)Hacker News / imp 10 / dev 20
- UEFA's Champions League draw creates unfair clusters; a Cayley graph fixes itHacker News / imp 20 / dev 60
- I trained a small transformer in 1.5hrs and it beats many LLMsHacker News / imp 55 / dev 75
- Quill (YC W20) Is Hiring a Fullstack SWEHacker News / imp 5 / dev 15
- Movie Scene Map – 13,312 films, series, games, anime and mangaHacker News / imp 20 / dev 25
- Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/sHacker News / imp 55 / dev 80
- Io_uring Without ReadaheadHacker News / imp 45 / dev 85
- We are rebuilding MonicaHacker News / imp 35 / dev 40
- FastpotifyHacker News / imp 20 / dev 50
- Introducing Ad Blocker for Firefox on iOSHacker News / imp 30 / dev 45
- A browser-based viewer for Office Open XML documentsHacker News / imp 40 / dev 70
- Tmp.0ut Volume 5Hacker News / imp 35 / dev 80
- Playa PhoneHacker News / imp 15 / dev 20
- I turned my security cameras into an automatic bird identification systemHacker News / imp 35 / dev 65
- Developing Enterprise Frontier Safeguards with our customersAnthropic News / imp 65 / dev 70
- Improving our alignment and security effortsAnthropic News / imp 75 / dev 80
- v2.1.257Claude Code / imp 60 / dev 85
- v2.1.252Claude Code / imp 40 / dev 75
- How AI-native companies turn workflows into operating capabilityOpenAI News / imp 45 / dev 55
- OpenAI supports California’s bill to advance youth AI safetyOpenAI News / imp 35 / dev 15
- Polimill builds Japan's next-generation public AI infrastructureOpenAI News / imp 40 / dev 50
- A milestone in expanding access to AIOpenAI News / imp 45 / dev 20
- Mapping global methane emissions from space with deep learningGoogle Research Blog / imp 50 / dev 65
- TimesFM-3: A zero-shot foundation model for multivariate forecastingGoogle Research Blog / imp 60 / dev 80
- Introducing agentic video understanding with GeminiGoogle DeepMind Blog / imp 55 / dev 75
- GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation modelsMicrosoft Research / imp 50 / dev 75
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AIHugging Face Blog / imp 55 / dev 85
- How we could save petabytes of cache storage with Zstandard and PingoraCloudflare Blog / imp 45 / dev 80
- Introducing Adaptive Intelligence: Undermining the economics of every bot attackCloudflare Blog / imp 50 / dev 75
- The Grails Plugin Has a New Home: Apache GrailsJetBrains Blog / imp 25 / dev 50
- Ensuring Code Compliance in Public Sector Software ProjectsJetBrains Blog / imp 35 / dev 60
- Authenticating TeamCity Builds to External Services With OIDCJetBrains Blog / imp 45 / dev 75
- Fine-Tuning SOTA Object Detection Models on Real-World DatasetsJetBrains Blog / imp 45 / dev 80
- From Leaderboards to Model Profiles: A Deep Dive Evaluation of LLMs for Agentic CodingJetBrains Blog / imp 70 / dev 85
- Sunsetting of the JetBrains Teacher Pack for BootcampsJetBrains Blog / imp 25 / dev 30
- Secure by default is your only way forwardDocker Blog / imp 50 / dev 70
- Building an Adaptive Agentic Cybersecurity System with NVIDIA NemotronNVIDIA Developer Blog / imp 55 / dev 80
- How to Size GPUs for AI Inference and TCO Without OverspendingNVIDIA Developer Blog / imp 45 / dev 75
- Run NVIDIA BioNeMo NIM Microservices for Protein Structure Prediction in Claude ScienceNVIDIA Developer Blog / imp 50 / dev 75
- Maximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPSNVIDIA Developer Blog / imp 45 / dev 80
- How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera RubinNVIDIA Developer Blog / imp 50 / dev 75
- The Hugging Face hack could indicate cultural issues at OpenAIMIT Technology Review (AI) / imp 75 / dev 85
- Pocket's AI made my game ideas real. Now Meta controls the results.Ars Technica (AI) / imp 35 / dev 50
- Sequoia-incubated Empirik launches with $21M to predict outages before they happenTechCrunch (AI) / imp 50 / dev 70
- Amazon Alexa can now alert you when something new might tempt you to shopTechCrunch (AI) / imp 35 / dev 45
- AIR raises $50M to help companies vet the skills and add-ons AI agents useTechCrunch (AI) / imp 60 / dev 80
- Fambot introduces an ‘AI chief of staff’ for familiesTechCrunch (AI) / imp 40 / dev 55
- Harvard Law dropout raises $6M for Blue Voice to build a ‘Harvey for police officers’TechCrunch (AI) / imp 45 / dev 60
- Clipto uses AI to search terabytes of video and is now valued at $250MTechCrunch (AI) / imp 50 / dev 65
- Nvidia’s $3.5B MediaTek bet reveals its plan for tackling Big Tech’s AI chip buildoutTechCrunch (AI) / imp 60 / dev 55
- Meeting note-taker Circleback adds a free tier to attract more customersTechCrunch (AI) / imp 25 / dev 35
- The US is building barriers around drones and robots, but China has scale to get around themTechCrunch (AI) / imp 50 / dev 40
- The rise of AI ‘civilizations’ and the fall of corporate responsibilityThe Verge (AI) / imp 60 / dev 70
- John Deere launched an AI chatbot for farmersThe Verge (AI) / imp 40 / dev 55
- Nvidia’s controversial DLSS 5 arrives September 3rd and requires serious GPU horsepowerThe Verge (AI) / imp 40 / dev 65
- Debian won’t ban AI code from its Linux distributionThe Verge (AI) / imp 55 / dev 70
- New York Governor Kathy Hochul thinks AI should be ‘less evil’The Verge (AI) / imp 35 / dev 20
- Texas Governor Abbott blocks funding for more Flock camerasThe Verge (AI) / imp 45 / dev 30
- OpenClaw 2.0 Releases with Simplified Setup and Collaborative AgentsInfoQ (AI/ML/Data Eng) / imp 55 / dev 80
- InfoQ previews the September cohorts of its online certification programsInfoQ (AI/ML/Data Eng) / imp 15 / dev 30
- HCP Terraform Positions Itself as the Control Plane for AI-Driven InfrastructureInfoQ (AI/ML/Data Eng) / imp 60 / dev 80
- DoorDash’s Flux Runs 130,000 Engineering Tasks Through Cloud-Based AgentsInfoQ (AI/ML/Data Eng) / imp 0 / dev 85
- Podcast: Scott Jenson on Evolving Desktop OS, Local-First, & Agentic UXInfoQ (AI/ML/Data Eng) / imp 45 / dev 65
- Presentation: Running AI at the Edge: Running Real Workloads Directly in the BrowserInfoQ (AI/ML/Data Eng) / imp 50 / dev 80
- Foundry Model Router Expands from Two Regions to 28, Refreshing Its Model PoolInfoQ (AI/ML/Data Eng) / imp 50 / dev 75
- Cash in on the AI Boom by Renting Out Your Spare ComputeIEEE Spectrum (AI) / imp 40 / dev 60
- Google Deepmind's new chief says frontier AI leadership is the only thing that mattersThe Decoder / imp 45 / dev 40
- Google's election AI Overviews are opaque, rely on few sources, and sometimes take sidesThe Decoder / imp 55 / dev 65
- Runway's Solaris is an AI system that generates software interfaces in real timeThe Decoder / imp 55 / dev 80
- Google's AI search dropped its emergency-call advice over nationalities but still flags people from FacebookThe Decoder / imp 50 / dev 60
- Bank of England chief warns that inflated AI valuations and rising leverage could trigger the next financial crisisThe Decoder / imp 55 / dev 30
- OpenAI says its ChatGPT ad business hits a $1 billion annual run rateThe Decoder / imp 40 / dev 20
- China's CXMT makes its first HBM3E chips, closing the AI memory gapThe Decoder / imp 60 / dev 75
- OpenAI starts charging some customers only when its AI actually worksThe Decoder / imp 50 / dev 40
- LWiAI Podcast #255 - Gemini 3.7, Jalapeño, Qwen 3.8, DronesLast Week in AI / imp 55 / dev 75
- Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local ModelarXiv cs.AI / imp 55 / dev 80
- Rating the Raters: Rasch Measurement Theory for LLM EvaluationarXiv cs.AI / imp 50 / dev 75
- Not All Explanations Are Sought: Information-Seeking Psychology for Human-Centered XAIarXiv cs.AI / imp 45 / dev 70
- Retrieving Relations, Detecting Fallacies: A RAG Approach to Political Debate AnalysisarXiv cs.AI / imp 45 / dev 75
- LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and OrientationarXiv cs.AI / imp 45 / dev 70
- Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient FieldsarXiv cs.AI / imp 55 / dev 75
- Class-Based Heuristic Selection for Solving the Flying Block PuzzlearXiv cs.AI / imp 35 / dev 50
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosarXiv cs.AI / imp 65 / dev 75
- Effectiveness of IoT and Deep Learning for Detection and Severity Assessment of Postelectrotermes militaris in Tea PlantationsarXiv cs.AI / imp 20 / dev 35
- Context Localization for Generalized Level-Based Evaluation in Knowledge-Based SystemsarXiv cs.AI / imp 35 / dev 50
- CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health IntelligencearXiv cs.AI / imp 55 / dev 65
- Thinking Costs Tokens: When More Structure is Worth the PricearXiv cs.AI / imp 65 / dev 85
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningarXiv cs.AI / imp 75 / dev 85
- SETU: An Agentic Ecosystem for Multilingual, Persona-Aware Communication CoachingarXiv cs.AI / imp 55 / dev 65
- Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety ModeratorarXiv cs.AI / imp 60 / dev 75
- LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety GuardrailsarXiv cs.AI / imp 65 / dev 85
- Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs)arXiv cs.AI / imp 40 / dev 35
- If Agents Were Angels, No Governance Would Be Necessary: Out-of-Band Policy Enforcement at a Trusted Tool BoundaryarXiv cs.AI / imp 75 / dev 85
- A Framework for Object-Centric Predictive Monitoring of Collaborative ProcessesarXiv cs.AI / imp 45 / dev 55
- Agents for Everyone: A Workshop Framework for Building Agentic AI Capabilities in a Distributed Curation CommunityarXiv cs.AI / imp 50 / dev 65
- PCFBench: A Diagnostic Benchmark for Product Carbon Footprint EstimationarXiv cs.AI / imp 55 / dev 75
- Probing Perceptual Priors of MLLMs via Gibbs Sampling with Interpretable Generative ControlsarXiv cs.AI / imp 50 / dev 75
- Why Didn't It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language ModelsarXiv cs.AI / imp 65 / dev 85
- Credo: Reusable Declarative Primitives for Agentic WorkflowsarXiv cs.AI / imp 0 / dev 85
- ReToolSQL: Agentic Reinforcement Learning for Robust Text-to-SQLarXiv cs.AI / imp 65 / dev 85
- CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied ActionarXiv cs.AI / imp 65 / dev 85
- CURA: Certified Runtime Alarms for Computer-Use AgentsarXiv cs.AI / imp 75 / dev 85
- AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference DynamicsarXiv cs.AI / imp 60 / dev 75
- Evidential-Based Higher-Order Set Argumentation FrameworkarXiv cs.AI / imp 35 / dev 40
- RealSWE: A Compositional Evaluation of Coding Agents under Realistic User RequestsarXiv cs.AI / imp 75 / dev 85
- KLOD: Locality-Preserving Knowledge Editing via Non-Target Distribution PreservationarXiv cs.AI / imp 50 / dev 75
- An Empirical Evaluation of Cross-City POI Recommendation on a Large-Scale BenchmarkarXiv cs.AI / imp 35 / dev 55
- From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical DiagnosisarXiv cs.AI / imp 60 / dev 75
- SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language ModelsarXiv cs.AI / imp 45 / dev 75
- CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction TuningarXiv cs.AI / imp 50 / dev 75
- See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEsarXiv cs.AI / imp 65 / dev 85
- HyQuant: Hybrid-Precision Quantization for LLM AttentionarXiv cs.AI / imp 50 / dev 75
- Resource Constraints and Performance in Agentic AI SystemsarXiv cs.AI / imp 75 / dev 85
- Rubric-to-Code Credit Assignment for Reinforcement LearningarXiv cs.AI / imp 65 / dev 85
- AI Alignment through a Game-theoretic Lens: A SurveyarXiv cs.AI / imp 60 / dev 65
- From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical ReasoningarXiv cs.AI / imp 50 / dev 75
- A Deep Learning-Based Stacking Ensemble Framework for Turbofan Engine Remaining Useful Life PredictionarXiv cs.AI / imp 35 / dev 65
- CASTANET: Causality-Aware Spatio-Temporal Adversarial Network Using Traffic Incident EffectsarXiv cs.AI / imp 35 / dev 55
- Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval DefensearXiv cs.AI / imp 75 / dev 85
- The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMsarXiv cs.AI / imp 60 / dev 75
- When Teacher Guidance Misleads: Reward-Aligned On-Policy DistillationarXiv cs.AI / imp 55 / dev 75
- SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch ProbingarXiv cs.AI / imp 65 / dev 85
- AERA: Adaptive Evidence Residual Allocation for Efficient Test-Time ReasoningarXiv cs.AI / imp 65 / dev 85
- openJiuwen: Beyond Static Harnesses for Long-Horizon Coding AgentsarXiv cs.AI / imp 0 / dev 90
- Learning from Hard Prompts: Difficulty-aware Advantage Amplification in Dynamic SamplingarXiv cs.AI / imp 45 / dev 75
- When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent SystemsarXiv cs.AI / imp 65 / dev 85
- GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent SocietiesarXiv cs.AI / imp 75 / dev 85
- Should I Use This Synthetic Dataset for Training? How to Test with Minimal Real DataarXiv cs.AI / imp 50 / dev 75
- Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language ModelarXiv cs.AI / imp 55 / dev 65
- PhenoIntel: A Lifecycle-Aligned Multi-Agent Web Application for Verified, Accessible Plant Phenotype AnalysisarXiv cs.AI / imp 55 / dev 75
- Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM AgentsarXiv cs.AI / imp 55 / dev 85
- String: An Agentic OS Where Every App Is a Markdown FilearXiv cs.AI / imp 75 / dev 95
- WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search AgentsarXiv cs.AI / imp 65 / dev 85
- Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement LearningarXiv cs.AI / imp 35 / dev 55
- SEPO: Evidence-Grounded Prompt Optimization via Structural EditingarXiv cs.AI / imp 65 / dev 85
- Speculative Probing: LLM Monitoring at Speculative-Decoding CostarXiv cs.AI / imp 65 / dev 85
- The Shape of Power: A Multilingual Framework for Social Power Reasoning in DialoguesarXiv cs.AI / imp 40 / dev 55
- Under-Mattress Temporal Sensing for Next-Day Agitation Risk Scoring in Dementia WardsarXiv cs.AI / imp 20 / dev 35
- CrabOS: An Operating System for Human-AI Co-inhabitationarXiv cs.AI / imp 75 / dev 85
- Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic EmbeddingsarXiv cs.AI / imp 45 / dev 65
- Generative AI Alignment with Hinduism's Theological Plurality and Sacred RepresentationarXiv cs.AI / imp 40 / dev 35
- Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar CompliancearXiv cs.AI / imp 65 / dev 85
- REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder FeaturesarXiv cs.AI / imp 65 / dev 85
- Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual EvaluationarXiv cs.AI / imp 75 / dev 85
- Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert SwitchingarXiv cs.AI / imp 55 / dev 75
- Physics-Guided Flow Matching for CT Image ReconstructionarXiv cs.AI / imp 45 / dev 65
- Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent CollaborationarXiv cs.AI / imp 65 / dev 85
- RECAST: Recent & Context-Aware Sampling for Test-Time Adaptation in Streaming BiosignalsarXiv cs.AI / imp 45 / dev 65
- LoopArena: Benchmarking Models as Runtime Controllers for Loop EngineeringarXiv cs.AI / imp 80 / dev 95
- Memristive-Friendly Hadamard Reservoir Computing: Structured, Multiplier-Free Recurrences at ScalearXiv cs.AI / imp 35 / dev 65
- MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in ChemistryarXiv cs.AI / imp 65 / dev 75
- Real-Valued Hyperdimensional Sequence Representations with Hadamard Product Binding and Shift EquivariancearXiv cs.AI / imp 35 / dev 65
- AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI AgentsarXiv cs.AI / imp 60 / dev 75
- Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelinesarXiv cs.AI / imp 55 / dev 75
- GRACE:Gradient-guided Coreset Selection for LLM UnlearningarXiv cs.AI / imp 50 / dev 75
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent HarnessesarXiv cs.AI / imp 80 / dev 95
- MAP: A Benchmark on Multimodal Accessibility Planning for Real World PlacesarXiv cs.AI / imp 55 / dev 65
- Timing-Aware Repurchase Prediction for Web-Scale E-Commerce: Survival Models for Multi-Surface Grocery RecommendationarXiv cs.AI / imp 40 / dev 65
- RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading AgentsarXiv cs.AI / imp 55 / dev 75
- VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC FilingsarXiv cs.AI / imp 55 / dev 75
- Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMsarXiv cs.AI / imp 75 / dev 95
- Prove2Me: An Open Collaborative Platform for Scaling Math FormalizationarXiv cs.AI / imp 65 / dev 85
- Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical ReasoningarXiv cs.AI / imp 65 / dev 85
- COVER: Identifiable Evaluation of Coalition RoutingarXiv cs.AI / imp 55 / dev 75
- AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video PredictionarXiv cs.AI / imp 45 / dev 65
- Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-ConfigurationarXiv cs.AI / imp 50 / dev 75
- When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AIarXiv cs.AI / imp 60 / dev 75
- InstructMesh: Selective Refinement of Generative 3D Models for FabricationarXiv cs.AI / imp 45 / dev 60
- Logos: An Agent Harness on a Cross-Process BusarXiv cs.AI / imp 85 / dev 95
- SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive InteractionarXiv cs.AI / imp 55 / dev 75
- Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate SpeecharXiv cs.AI / imp 45 / dev 65
- The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial ModelsarXiv cs.AI / imp 55 / dev 65
- PACE: Publisher-Adaptive Content Extraction via Agentic AutomationarXiv cs.AI / imp 65 / dev 85
- UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Question AnsweringarXiv cs.AI / imp 45 / dev 65
- Select, Don't Train: The Benefits of Modular Entity Disambiguation with LLM-Based SelectionarXiv cs.AI / imp 55 / dev 75
- XHotpotQA: A Benchmark for Cross-Lingual Knowledge Composition in Multi-Hop Question AnsweringarXiv cs.AI / imp 55 / dev 75
- A Survey on Rubric-Guided Reinforcement Learning for Language ModelsarXiv cs.AI / imp 65 / dev 75
- Marginal Coverage Credit Reduces Redundant Exploration in Parallel State-Entropy OptimizationarXiv cs.AI / imp 55 / dev 85
- Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment GaparXiv cs.AI / imp 65 / dev 85
- DAMP: Decay-Aware Mixed-Precision Recurrent-State QuantizationarXiv cs.AI / imp 35 / dev 75
- Trajectory-Level Speculative Decoding for Diffusion Language ModelsarXiv cs.AI / imp 30 / dev 70
- Destroy Me: Automatic Artifact Generation for Histopathology ImagesarXiv cs.AI / imp 15 / dev 50
- FVeinSyn: Synthetic Finger Vein Image GeneratorarXiv cs.AI / imp 10 / dev 45
- Self-Explainable Multi-Label Graph Neural Network for Correlated Evidence AttributionarXiv cs.AI / imp 20 / dev 60
- Quanta Perception as Probabilistic EventsarXiv cs.AI / imp 25 / dev 40
- PHR-VLA: Planning Horizon Reasoning for Vision-Language-Action ModelsarXiv cs.AI / imp 35 / dev 65
- Tensor-Accelerated Eager Multi-Resolution Grids for Evolving Large-Scale SubstratesarXiv cs.AI / imp 8 / dev 50
- LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State DataarXiv cs.AI / imp 30 / dev 60
- Depth-Aware Pothole Detection Using YOLO and RT-DETR at the EdgearXiv cs.AI / imp 12 / dev 45
- Curvature-Aware Radius Shrinkage for Adaptive Nearest Neighbor ClassificationarXiv cs.AI / imp 8 / dev 55
- Knowing Before Answering: Decoding Language Models for Reliable RAGarXiv cs.AI / imp 45 / dev 70
- Semantic Watermarking with Order-Robust Detection over Sub-sentence UnitsarXiv cs.AI / imp 35 / dev 65
- First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game AgentsarXiv cs.AI / imp 18 / dev 50
- CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CTarXiv cs.AI / imp 20 / dev 45
- Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter DistancearXiv cs.AI / imp 12 / dev 60
- RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression TestingarXiv cs.AI / imp 40 / dev 70
- Efficient Auto-Interpretability of AI Models in BiologyarXiv cs.AI / imp 35 / dev 65
- Beyond Search-Imitation: Prior-Directed Exploration for Searchless ChessarXiv cs.AI / imp 15 / dev 55
- Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal ConflictarXiv cs.AI / imp 28 / dev 65
- How Much Can AI Understand? Toward AI-Assisted Sensemaking of Collaborative Discussion in Groups with Shared HistoryarXiv cs.AI / imp 25 / dev 50
- ContextLeak: Exfiltrating LLM Agent Context via Malicious ToolsarXiv cs.AI / imp 60 / dev 75
- Actionable CBFI: Integrating Structural Decomposition and Causal Counterfactual Recourse for Tabular Machine LearningarXiv cs.AI / imp 22 / dev 60
- FISGuard: Defending Against Membership Inference via Fixed Input SubspacesarXiv cs.AI / imp 35 / dev 65
- FedEHR-Agents: Federated Agentic Optimization for Automated EHR ModelingarXiv cs.AI / imp 40 / dev 70
- From Perspective to Fisheye Depth Estimation and Open-Vocabulary SegmentationarXiv cs.AI / imp 20 / dev 60
- SOMTab: Set-Order Mamba for Efficient Tabular In-Context LearningarXiv cs.AI / imp 50 / dev 75
- OpenStamp: A Watermark for Open-Source Language ModelsarXiv cs.AI / imp 40 / dev 70
- LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing PagesarXiv cs.AI / imp 25 / dev 60
- Low-Altitude Fluid Antenna Network with Multi-Agent Reinforcement LearningarXiv cs.AI / imp 20 / dev 55
- PCBnet: A Dataset and Automatic Construction of SPICE Netlists from Schematic ImagesarXiv cs.AI / imp 25 / dev 60
- Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering ResearchersarXiv cs.AI / imp 35 / dev 50
- Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM VerificationarXiv cs.AI / imp 40 / dev 70
- CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?arXiv cs.AI / imp 65 / dev 75
- A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation ConstraintarXiv cs.AI / imp 25 / dev 70
- When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?arXiv cs.AI / imp 18 / dev 65
- Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded ReasoningarXiv cs.AI / imp 45 / dev 70
- Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-CodearXiv cs.AI / imp 55 / dev 75
- SimpCue: Cue-Based Prompting for Multilingual Text SimplificationarXiv cs.AI / imp 20 / dev 55
- Explainable Uncertainty Estimation for Reliable Medical AIarXiv cs.AI / imp 35 / dev 65
- Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language ModelsarXiv cs.AI / imp 50 / dev 70
- VersaGauss: A Versatile Framework for Generating Multiphase Dynamics with 3D GaussiansarXiv cs.AI / imp 15 / dev 60
- Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual RepresentationsarXiv cs.AI / imp 28 / dev 65
- VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement LearningarXiv cs.AI / imp 55 / dev 75
- CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest RadiographsarXiv cs.AI / imp 20 / dev 55
- Post-Edit Re-Verification in Simulator-Backed Engineering Agents: A Controlled Comparison of Verification-Cadence GuidancearXiv cs.AI / imp 50 / dev 70
- The Approximation Rank of Softmax Attention: Sharp Geometric Laws and Robust Interaction DimensionarXiv cs.AI / imp 15 / dev 60
- Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative ResultarXiv cs.AI / imp 22 / dev 65
- Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous SystemsarXiv cs.AI / imp 30 / dev 70
- Text Restoration of Ancient Documents with Language ModelsarXiv cs.AI / imp 20 / dev 50
- Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk ControlarXiv cs.AI / imp 20 / dev 55
- Beyond Flat Netlist: Hierarchical Graph Representation Learning for Scalable Analysis of Sequential CircuitsarXiv cs.AI / imp 30 / dev 70
- Performative Privacy: When Differential Privacy Maximizes UtilityarXiv cs.AI / imp 35 / dev 65
- Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface ScoringarXiv cs.AI / imp 22 / dev 60
- A comprehensive and trustworthy benchmark of AI methods for change detection in Earth observationarXiv cs.AI / imp 25 / dev 60
- Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search OperationsarXiv cs.AI / imp 30 / dev 65
- Embedding Models for Stance-Aware Argument RetrievalarXiv cs.AI / imp 18 / dev 60
- A Probabilistic Interpretation of KV Cache EvictionarXiv cs.AI / imp 55 / dev 75
- MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel OperationarXiv cs.AI / imp 40 / dev 70
- PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel OperationarXiv cs.AI / imp 40 / dev 70
- VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-DistillationarXiv cs.AI / imp 25 / dev 65
- Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and LossarXiv cs.AI / imp 35 / dev 70
- Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense LayersarXiv cs.AI / imp 50 / dev 75
- BanglaMed-QA: A Question Answering System for Healthcare Support in BanglaarXiv cs.AI / imp 18 / dev 55
- Cross-Spectral Dense Correspondence for Multimodal Spectral Medical ImagingarXiv cs.AI / imp 18 / dev 55
- Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test TakersarXiv cs.AI / imp 22 / dev 55
- Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility PilotarXiv cs.AI / imp 15 / dev 45
- AI as Teammate: Rethinking Task Distribution in Medical TrainingarXiv cs.AI / imp 28 / dev 45
- When Linguistic and Internal Confidence Diverge in Large Language ModelsarXiv cs.AI / imp 40 / dev 70
- LongPIBench: A Long-Context Benchmark for Prompt InjectionarXiv cs.AI / imp 55 / dev 75
- Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQLarXiv cs.AI / imp 50 / dev 75
- Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet ExtractionarXiv cs.AI / imp 45 / dev 70
- ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CTarXiv cs.AI / imp 25 / dev 65
- Anatomy-Aware Promptable Segmentation with Online Interactive Training for AUTOPET VarXiv cs.AI / imp 22 / dev 60
- Real-time virtual circuits for plasma shape control via neural network emulators: experimental demonstration on MAST UpgradearXiv cs.AI / imp 20 / dev 55
- NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometryarXiv cs.AI / imp 30 / dev 70
- How Proper Scoring Rules Shape LLM ForecastingarXiv cs.AI / imp 35 / dev 70
- LLM-Based Agents for Software and Systems Security: Approaches, Applications, and AssessmentarXiv cs.AI / imp 60 / dev 75
- On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin MarketplacesarXiv cs.AI / imp 65 / dev 75
- Conformal Uncertainty Quantification Guarantees for Neural OperatorsarXiv cs.AI / imp 20 / dev 65
- Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural NetworksarXiv cs.AI / imp 8 / dev 50
- An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World ModelsarXiv cs.AI / imp 18 / dev 60
- Video Generative Models as Geometry LearnerarXiv cs.AI / imp 22 / dev 65
- Blog: Survey of OptimizersarXiv cs.AI / imp 45 / dev 75
- Learning a Size-Weight Frontier for Synthetic-Augmented InferencearXiv cs.AI / imp 28 / dev 65
- Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation LearningarXiv cs.AI / imp 20 / dev 50
- Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes ReasoningarXiv cs.AI / imp 35 / dev 70
- BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment UnderstandingarXiv cs.AI / imp 25 / dev 60
- Multimodal Collaborative Debate for Zero-Shot Time Series ReasoningarXiv cs.AI / imp 45 / dev 70
- Real-Time AI Service Economy: A Framework for Agentic Computing Across the ContinuumarXiv cs.AI / imp 65 / dev 80
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World ModelsarXiv cs.AI / imp 55 / dev 75
- PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage NormalizationarXiv cs.AI / imp 55 / dev 75
- Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model PredictionsarXiv cs.AI / imp 40 / dev 70
- Understanding and Enforcing Weight Disentanglement in Task ArithmeticarXiv cs.AI / imp 32 / dev 70
- D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven DiscoveryarXiv cs.AI / imp 60 / dev 75
- Rethinking Vacuity for OOD Detection in Evidential Deep LearningarXiv cs.AI / imp 22 / dev 60
- Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan GenerationarXiv cs.AI / imp 30 / dev 65
- ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMsarXiv cs.AI / imp 55 / dev 75
- AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical PropertiesarXiv cs.AI / imp 35 / dev 70
- RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse EvaluationarXiv cs.AI / imp 28 / dev 65
- Flow Reasoning Models: Turning Discrete Flows Into Efficient Recurrent ReasonersarXiv cs.AI / imp 50 / dev 75
- APeB: Benchmarking Personalization Ability of Large Language Model AgentsarXiv cs.AI / imp 45 / dev 70
- Atomic Units of X: The Compression Layer of IntelligencearXiv cs.AI / imp 30 / dev 50
- Set-shifting Behavioral Test for Harnessed AgentsarXiv cs.AI / imp 65 / dev 80
- SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question AnsweringarXiv cs.AI / imp 50 / dev 75
- HALT: Verification-Aware Stopping for Retrieval-Augmented Search AgentsarXiv cs.AI / imp 55 / dev 80
- Agentao: A Policy-Governed Runtime Harness for Embeddable Tool-Using LLM AgentsarXiv cs.AI / imp 75 / dev 85
- When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit SeparationarXiv cs.AI / imp 50 / dev 70
- RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL TrainingarXiv cs.AI / imp 60 / dev 85
- When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content CreationarXiv cs.AI / imp 35 / dev 50
- STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic ControlarXiv cs.AI / imp 65 / dev 85
- Does Rank Still Matter? Position Bias When AI Agents Shop on Our BehalfarXiv cs.AI / imp 55 / dev 65
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering VectorsarXiv cs.AI / imp 75 / dev 85
- Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval GroundingarXiv cs.AI / imp 60 / dev 80
- SKILL.state: Scalable Long-Horizon Agent SkillsarXiv cs.AI / imp 75 / dev 85
- AgentFold: Closed-Loop Agentic Search for Protein Folding Model DesignarXiv cs.AI / imp 70 / dev 80
- Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron OccupationarXiv cs.AI / imp 50 / dev 65
- Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A SurveyarXiv cs.AI / imp 55 / dev 75
- Evaluating the Performance of Large Language Models on GAOKAO BenchmarkarXiv cs.AI / imp 30 / dev 60
- Let the Flows Tell: Solving Graph Combinatorial Optimization Problems with GFlowNetsarXiv cs.AI / imp 50 / dev 75
- Long Story Short: Story-level Video Understanding from 20K Short FilmsarXiv cs.AI / imp 40 / dev 60
- PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data SelectionarXiv cs.AI / imp 50 / dev 75
- Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social SituationsarXiv cs.AI / imp 45 / dev 70
- Attention as Conditioning: What Classical Learning Theory Predicts About Linear TransformersarXiv cs.AI / imp 35 / dev 70
- Beyond the Rosetta Stone: Unification Forces in Generalization DynamicsarXiv cs.AI / imp 40 / dev 65
- Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep LearningarXiv cs.AI / imp 35 / dev 55
- Steering Multimodal Large Language Models Decoding for Context-Aware SafetyarXiv cs.AI / imp 70 / dev 80
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language ModelsarXiv cs.AI / imp 45 / dev 70
- Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in Language-based Object DetectionarXiv cs.AI / imp 50 / dev 75
- OceanGym: A Benchmark Environment for Underwater Embodied AgentsarXiv cs.AI / imp 50 / dev 75
- PRISM: Agentic Retrieval with LLMs for Multi-Hop Question AnsweringarXiv cs.AI / imp 65 / dev 85
- Riverbank Erosion Analysis in Bangladesh Using Spatiotemporal SegmentationarXiv cs.AI / imp 30 / dev 50
- Quantifying Affective Bias in Low-Resource Media: Large-Scale Emotion Profiling of Bengali HeadlinesarXiv cs.AI / imp 35 / dev 50
- Think-at-Hard: Dynamic Looped Transformers for Improved ReasoningarXiv cs.AI / imp 60 / dev 80
- OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular FusionarXiv cs.AI / imp 50 / dev 75
- The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal BehaviorarXiv cs.AI / imp 75 / dev 80
- FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech AdaptationarXiv cs.AI / imp 55 / dev 80
- Aligning Agentic World Models via Knowledgeable Experience LearningarXiv cs.AI / imp 70 / dev 85
- CoFrGeNet: Continued Fraction Architectures for Language GenerationarXiv cs.AI / imp 40 / dev 70
- Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic ReasoningarXiv cs.AI / imp 50 / dev 70
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsarXiv cs.AI / imp 65 / dev 85
- ASA: Backbone-Training-Free Representation Engineering for Tool-Calling AgentsarXiv cs.AI / imp 75 / dev 85
- FENCE: A Financial and Multimodal Jailbreak Detection DatasetarXiv cs.AI / imp 70 / dev 80
- From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to ThemselvesarXiv cs.AI / imp 75 / dev 85
- Large Reasoning Models Struggle to Transfer Parametric Knowledge Across ScriptsarXiv cs.AI / imp 50 / dev 75
- InfoMamba: An Attention-Free Hybrid Mamba-Transformer ModelarXiv cs.AI / imp 60 / dev 80
- The Autonomy Tax: Defense Training Breaks LLM AgentsarXiv cs.AI / imp 80 / dev 85
- Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture - Bridging Predictive and Generative Self-Supervised LearningarXiv cs.AI / imp 50 / dev 75
- Select, Label, Evaluate: Active Testing in NLParXiv cs.AI / imp 50 / dev 75
- Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta EvidencearXiv cs.AI / imp 50 / dev 70
- Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement LearningarXiv cs.AI / imp 60 / dev 80
- PolicyLong: Towards On-Policy Context ExtensionarXiv cs.AI / imp 65 / dev 85
- Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding TasksarXiv cs.AI / imp 65 / dev 85
- Benefits of Low-Cost Bio-Inspiration in the Age of OverparametrizationarXiv cs.AI / imp 40 / dev 65
- Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMsarXiv cs.AI / imp 45 / dev 60
- G-Loss: Graph-Guided Fine-Tuning of Language ModelsarXiv cs.AI / imp 50 / dev 80
- ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and SpacearXiv cs.AI / imp 50 / dev 75
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack SurfacesarXiv cs.AI / imp 80 / dev 85
- Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access ControlarXiv cs.AI / imp 85 / dev 90
- SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development GoalsarXiv cs.AI / imp 45 / dev 65
- More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of ActivationsarXiv cs.AI / imp 55 / dev 80
- Negligible in Size, Significant in Effect: On Scale Vectors in Large Language ModelsarXiv cs.AI / imp 50 / dev 75
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data AnalysisarXiv cs.AI / imp 70 / dev 85
- DiffuSent: Towards a Unified Diffusion Framework for Aspect-Based Sentiment AnalysisarXiv cs.AI / imp 45 / dev 70
- The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini ModelsarXiv cs.AI / imp 70 / dev 75
- TokenPilot: Cache-Efficient Context Management for LLM AgentsarXiv cs.AI / imp 75 / dev 85
- The Discrete-Log Clock: How a Transformer Learns Modular MultiplicationarXiv cs.AI / imp 35 / dev 70
- CASPER in the Machine: Insights into Character Variety in LLM-Generated StoriesarXiv cs.AI / imp 40 / dev 55
- An LLM-Based Framework for Intent-Driven Network Topology DesignarXiv cs.AI / imp 55 / dev 75
- GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid ReasoningarXiv cs.AI / imp 50 / dev 70
- On the Depth Scalability of Logic Gate NetworksarXiv cs.AI / imp 40 / dev 70
- REPREC: Representation Driven Parameter-Efficient Recommendation SystemarXiv cs.AI / imp 50 / dev 75
- Where Steering Signals Come From: Activation Source Selection in Activation SteeringarXiv cs.AI / imp 55 / dev 80
- Locked Evaluation Surfaces: Transfer Failure and Sampling-Depth Entanglement in CRISPRi Perturbation-Effect PredictionarXiv cs.AI / imp 40 / dev 65
- Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search AgentsarXiv cs.AI / imp 70 / dev 85
- ED-CSP: Crystal Structure Prediction from Electron DiffractionarXiv cs.AI / imp 45 / dev 70
- BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers InferencearXiv cs.AI / imp 60 / dev 80
- RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language NavigationarXiv cs.AI / imp 55 / dev 80
- PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal ModelsarXiv cs.AI / imp 45 / dev 70
- How Far Should Tokenization Go? Predictive Effectiveness and Relational LosslessnessarXiv cs.AI / imp 55 / dev 85
- JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded VerificationarXiv cs.AI / imp 70 / dev 85
- Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented GenerationarXiv cs.AI / imp 75 / dev 85
- Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic ConstraintsarXiv cs.AI / imp 70 / dev 85
- GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-NetsarXiv cs.AI / imp 45 / dev 75
- Multi-Winner Voting with Argumentative BallotsarXiv cs.AI / imp 35 / dev 55
- Macro-Operator Generation and Predicate Selection for TAMP Operator LearningarXiv cs.AI / imp 60 / dev 80
- On-policy Distillation with Verifiable RewardarXiv cs.AI / imp 70 / dev 85
- SpecMine: A Large-Scale Corpus of Spec-Driven Development ArtifactsarXiv cs.AI / imp 0 / dev 95
- MathAdv: What Theorem Provers Know, Reason, Formalize, and GeneralizearXiv cs.AI / imp 65 / dev 85
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent MemoryarXiv cs.AI / imp 75 / dev 85
- TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning DevelopmentarXiv cs.AI / imp 75 / dev 90
- AI Models Can Predict and Collaboratively Modulate Human Memory SearcharXiv cs.AI / imp 60 / dev 70
- Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM EvaluationarXiv cs.AI / imp 75 / dev 80
- Comparing Chunking and Embedding Strategies for Turkish RAG SystemsarXiv cs.AI / imp 60 / dev 80
- Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AIarXiv cs.AI / imp 85 / dev 95
- LiveVVT: High-Fidelity Video Virtual Try-On in Real TimearXiv cs.AI / imp 50 / dev 75
- Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM AgentsarXiv cs.AI / imp 85 / dev 90
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?arXiv cs.AI / imp 60 / dev 80
- ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM ReasoningarXiv cs.LG / imp 70 / dev 85
- Unsupervised Latent Space Alignment with Hyperspherical Geodesic MatchingarXiv cs.LG / imp 50 / dev 75
- Curvature Cryptanalysis of Smooth Transformer Feed-Forward NetworksarXiv cs.LG / imp 35 / dev 45
- Equivariant Sheaf Neural Networks: Learning Geometric Transport on GraphsarXiv cs.LG / imp 25 / dev 35
- The Halt Vector: Internalizing a Causal Steering Intervention for Efficient ReasoningarXiv cs.LG / imp 60 / dev 60
- Conservative Hybrid Graph Networks for Process Systems with Learned RoutingarXiv cs.LG / imp 30 / dev 40
- Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?arXiv cs.LG / imp 35 / dev 50
- Learning-Theoretic Foundation for General Coded Computing: The Straggler SettingarXiv cs.LG / imp 40 / dev 50
- SemKV: Semantic Mixed-Precision KV Cache Quantization Guided by the Quality Cliff for Long-Context LLM InferencearXiv cs.LG / imp 75 / dev 75
- RankShift: In-Database Detection and Explanation of Categorical ShiftsarXiv cs.LG / imp 30 / dev 40
- Revisiting the Provable-Auditable Privacy Gap of DP-SGDarXiv cs.LG / imp 45 / dev 50
- From the Loss Landscape to Diverse Feature Learning in Neural NetworksarXiv cs.LG / imp 35 / dev 40
- Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVIarXiv cs.LG / imp 25 / dev 35
- V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational AwarenessarXiv cs.LG / imp 25 / dev 30
- Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia DataarXiv cs.LG / imp 30 / dev 40
- Context-Aware Interpretable Representations for Retrieval and Graph Convolutional Network ClassificationarXiv cs.LG / imp 35 / dev 45
- Hybrid Semantic Context-Enhanced Ensemble Learning for Wind Power Ramp-Event Forecasting and Uncertainty-Aware EvaluationarXiv cs.LG / imp 30 / dev 40
- Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World ModelsarXiv cs.LG / imp 50 / dev 50
- NVE: A Separability and Coverage-Aware Internal Validation Metric for BiclusteringarXiv cs.LG / imp 20 / dev 30
- Sparse Koopman Autoencoders Identify Local Dynamical Regimes in Multibasin SystemsarXiv cs.LG / imp 25 / dev 35
- PathBridger: Subgoal Bridges for Offline Goal-Conditioned Reinforcement LearningarXiv cs.LG / imp 45 / dev 50
- Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and SteeringarXiv cs.LG / imp 60 / dev 60
- Titans-QFWP: A Regime-Aware Hybrid Quantum Fast Weight Programmer for Portfolio OptimizationarXiv cs.LG / imp 15 / dev 20
- Development of an Autonomous AI Coding Agent using Monte Carlo Tree Search (MCTS) and Gemini LLM FrameworksarXiv cs.LG / imp 70 / dev 80
- Temperature-Adaptive Transformed Teacher MatchingarXiv cs.LG / imp 40 / dev 55
- PathGuide: Dynamic Classifier-Free Guidance via On-Policy Transport AlignmentarXiv cs.LG / imp 45 / dev 50
- Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor ProfilingarXiv cs.LG / imp 35 / dev 50
- Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution SpacearXiv cs.LG / imp 50 / dev 55
- HalluPrism: When Multimodal Uncertainty Should Diagnose, Not DecidearXiv cs.LG / imp 55 / dev 60
- PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGCarXiv cs.LG / imp 30 / dev 40
- RL-FAT: Reinforcement Learning for Fair Adversarial TrainingarXiv cs.LG / imp 45 / dev 55
- Adaptive Multi-Branching for Shallow Decision Tree InductionarXiv cs.LG / imp 35 / dev 45
- When Do Larger Batches Help Scale LLM Reinforcement Learning?arXiv cs.LG / imp 55 / dev 65
- A Spectral Identifiability Threshold for Dissipative Rate Recovery from Truncated Liouvillian SpectraarXiv cs.LG / imp 15 / dev 20
- MEL: Coordinate-Preserving EEG Tokenization for fMRI TranslationarXiv cs.LG / imp 30 / dev 35
- Information-Based Calibration of Uncertainty Quantification in Product-of-Experts Gaussian Process ModelsarXiv cs.LG / imp 40 / dev 50
- Spatial Entropy based Partitioning for Spatiotemporal Graph UnlearningarXiv cs.LG / imp 45 / dev 55
- Unlearning on Spatio-Temporal Graphs through Subgraph Virtual Edge ReconstructionarXiv cs.LG / imp 45 / dev 55
- Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on MultipliersarXiv cs.LG / imp 40 / dev 50
- Where Induction Runs Out: Description-Length Difficulty and the Memorisation Gap in Integer-Sequence BenchmarksarXiv cs.LG / imp 45 / dev 55
- Scalable Clinical Data Infrastructure and Comparative ML Evaluation for Hospitalisation Risk Prediction in Elderly Patients with Multiple Long-Term Conditions using CPRDarXiv cs.LG / imp 50 / dev 60
- One Capability or Many? Testing the Economic Validity of Frontier AI EvaluationarXiv cs.LG / imp 65 / dev 70
- Behavioral Latency as Weak Event-Time Supervision for EEG Reaction-Time DecodingarXiv cs.LG / imp 25 / dev 35
- Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric ObservationsarXiv cs.LG / imp 50 / dev 55
- SS-ESOAP: Self-Scaled Adaptive Preconditioning for Physics-Informed LearningarXiv cs.LG / imp 40 / dev 55
- Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged CapabilitiesarXiv cs.LG / imp 55 / dev 60
- A Causal Model for Locating and Unlocking Sandbagging in Model OrganismsarXiv cs.LG / imp 60 / dev 65
- Knowledge Distillation under Teacher Misspecification: An Order-Parameter Analysis of the Gap between Teacher Mimicry and Task PerformancearXiv cs.LG / imp 40 / dev 55
- Learning Human Health and Diseases from 24-hour Wrist MovementarXiv cs.LG / imp 45 / dev 50
- Target-Aware State-Adaptive $p$-Dirichlet Graph Neural Regression for Non-Invasive Body-Composition EstimationarXiv cs.LG / imp 25 / dev 40
- Adversarial Online Classification with a PreviewarXiv cs.LG / imp 30 / dev 40
- Denoising as Projection: Constrained Optimization with Gradient-Guided DiffusionarXiv cs.LG / imp 45 / dev 55
- On the Plasticity Collapse in Continual Machine UnlearningarXiv cs.LG / imp 50 / dev 60
- MedCache: Efficient and Temporally Valid Memory for Longitudinal Clinical AgentsarXiv cs.LG / imp 60 / dev 70
- BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive LearningarXiv cs.LG / imp 40 / dev 50
- Which LLM for Which Work? Budgeted Model Allocation under Uncertain EvaluationarXiv cs.LG / imp 65 / dev 75
- Asynchronous Cooperative Online Learning for Multi-Robot Control under Computational DelaysarXiv cs.LG / imp 40 / dev 55
- HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession PlanningarXiv cs.LG / imp 30 / dev 40
- Event-triggered Control and Online Learning for Networked Systems under Computational DelaysarXiv cs.LG / imp 40 / dev 55
- Predicting the Unpredictable: LLM-powered Long-term Chaotic Time Series Forecasting under Short-term ObservationsarXiv cs.LG / imp 50 / dev 65
- On the Resilience of Text-to-Video Diffusion Models to Hardware FaultsarXiv cs.LG / imp 40 / dev 55
- Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User BehaviorarXiv cs.LG / imp 40 / dev 60
- Wide Learning: Learning to Reach EvidencearXiv cs.LG / imp 40 / dev 50
- Unsupervised Multi-Scale Gromov-Wasserstein Hypergraph AlignmentarXiv cs.LG / imp 30 / dev 40
- LLMODE: Aligning ODEs with LLMs via Gated Token Injection for Irregular Spatio-Temporal ForecastingarXiv cs.LG / imp 50 / dev 65
- Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient FlowarXiv cs.LG / imp 50 / dev 60
- A Target-Centric Survey of Quantization-Aware TrainingarXiv cs.LG / imp 60 / dev 75
- Creation begins with understanding: LLMs as strategy designers for privacy-preserving tabular data synthesisarXiv cs.LG / imp 50 / dev 65
- Last Step Matters: Early Uncertainty Cannot Predict Failure in Long-Horizon AgentsarXiv cs.LG / imp 60 / dev 70
- Higher-Dimensional Rotary Position EmbeddingarXiv cs.LG / imp 45 / dev 65
- GraM-Diff: A Unified Graph-Mamba Diffusion Framework for EEG-Based Alzheimer's Disease Data Generation and DiagnosisarXiv cs.LG / imp 40 / dev 50
- ECA-BLS: An Efficient Complex-Augmented Broad Learning SystemarXiv cs.LG / imp 30 / dev 45
- PruneShift: A Framework for Evaluating Decision Reliability in Structured PruningarXiv cs.LG / imp 40 / dev 60
- Structure Aware Neural Architecture Search for Mixture of ExpertsarXiv cs.LG / imp 55 / dev 70
- Designing for the Next Click: Bandits for Real-Time Page LayoutarXiv cs.LG / imp 45 / dev 60
- Uncertainty-Driven Replay Memory for Reinforcement LearningarXiv cs.LG / imp 45 / dev 60
- Partially Linear Autoencoders for Manifold Learning and Dimensionality ReductionarXiv cs.LG / imp 35 / dev 50
- Towards an Expressivity-Normalized Energy-Demand Comparison of ANNs and SNNsarXiv cs.LG / imp 35 / dev 50
- Structural Hierarchy and Geometry in Molecular Representation LearningarXiv cs.LG / imp 35 / dev 50
- Sensitivity-Constrained Neural Operators for Data-Efficient Forward and Inverse Modeling of Partial Differential Equation SystemsarXiv cs.LG / imp 40 / dev 60
- Joint Spatiotemporal Spectral Neural Operators for Learning PDEs on Irregular DomainsarXiv cs.LG / imp 45 / dev 65
- INTERVenE: Temporal-Abstraction-Interval Based Transformers for Short-Horizon Medical Event PredictionarXiv cs.LG / imp 55 / dev 70
- Diffusion-Based Inverse Design of Dielectric Resonator Metasurfaces for Shaping Smart Electromagnetic EnvironmentsarXiv cs.LG / imp 30 / dev 40
- On the Recoverability of Private Information Unlearning in Large Language ModelsarXiv cs.LG / imp 60 / dev 70
- Robust Broad Learning System with Wave Loss for Classification under Data UncertaintyarXiv cs.LG / imp 35 / dev 50
- The Intervention Gap in Latent World ModelsarXiv cs.LG / imp 55 / dev 65
- Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language ModelsarXiv cs.LG / imp 55 / dev 70
- Multiclass Linear Perceptrons with Multiplicative MarginsarXiv cs.LG / imp 30 / dev 45
- Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label CorrectionarXiv cs.LG / imp 50 / dev 65
- When 3D Gaussian Splatting Recovers Real SurfacesarXiv cs.LG / imp 45 / dev 60
- How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral AccountarXiv cs.LG / imp 70 / dev 80
- Selection, Representation, and Execution in Sparse Fourier Neural OperatorsarXiv cs.LG / imp 40 / dev 60
- Tracing Generated Samples to Training-Data Clusters in Flow-Matching ModelsarXiv cs.LG / imp 50 / dev 60
- A Lightweight Phenology-Aware YOLOv5 Framework for Tomato Growth Stage Detection in Resource-Constrained Bhutanese Greenhouse EnvironmentsarXiv cs.LG / imp 30 / dev 40
- SMOTE-VAR: An Uncertainty-Aware Oversampling Method for Predicting Depression Remission in University StudentsarXiv cs.LG / imp 45 / dev 60
- Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear OptimizationarXiv cs.LG / imp 40 / dev 60
- Supraglacial Lake Fate Is Knowable Long Before the Season EndsarXiv cs.LG / imp 30 / dev 40
- TPR-Attention for Combinatorial GeneralizationarXiv cs.LG / imp 55 / dev 70
- Converse and Collision-Based Achievability for Node Localization with Hybrid Distance-Spectral Graph Positional EncodingsarXiv cs.LG / imp 30 / dev 50
- Reinforcement Learning for Symbolic Equation SolvingarXiv cs.LG / imp 60 / dev 75
- Benchmarking Peptide-Protein Affinity Prediction Across Peptide and Target ShiftsarXiv cs.LG / imp 40 / dev 55
- Certified Safety Radii in Forecast-Error Space for Wasserstein Distributionally Robust Small Signal Stability-Constrained AC Optimal Power Flow via Lifted Spectrahedral ContainmentarXiv cs.LG / imp 35 / dev 50
- Diffusion-Based Refinement for Kilometer-Scale Probabilistic Precipitation NowcastingarXiv cs.LG / imp 20 / dev 50
- Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV CachearXiv cs.LG / imp 60 / dev 80
- Exact Recovery Thresholds for Weighted Data Selection in Vector-Valued Linear RegressionarXiv cs.LG / imp 10 / dev 30
- Multivariate Scientific Data Compression with Learned Cross-Variable Latent Decorrelation and Autoregressive Entropy ModelingarXiv cs.LG / imp 30 / dev 60
- BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement LearningarXiv cs.LG / imp 35 / dev 70
- CateKV: On Sequential Consistency for Long-Context LLM Inference AccelerationarXiv cs.LG / imp 60 / dev 80
- Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMsarXiv cs.LG / imp 55 / dev 80
- Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small InitializationarXiv cs.LG / imp 40 / dev 60
- Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation GuidancearXiv cs.LG / imp 25 / dev 50
- Generative multi-domain transfer learning for fault detection in data-scarce wind turbinesarXiv cs.LG / imp 30 / dev 55
- Learning PDE Time-Stepping with Neural Cellular AutomataarXiv cs.LG / imp 45 / dev 65
- Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profilingarXiv cs.LG / imp 35 / dev 50
- Beyond Churn: Predicting Financial Fragmentation in Retail Banking with Temporal Machine LearningarXiv cs.LG / imp 30 / dev 50
- Mode Connectivity Beyond Classifiers: Evidence from Generative and Contrastive ModelsarXiv cs.LG / imp 45 / dev 65
- Beat-Synchronous Tokenization for ECG TransformersarXiv cs.LG / imp 35 / dev 55
- Convergence rates for the RMSprop optimizer with full control of the hyperparametersarXiv cs.LG / imp 40 / dev 70
- RSLM: Training-Free Vector Quantization for Approximate Nearest Neighbor SearcharXiv cs.LG / imp 50 / dev 75
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention ServingarXiv cs.LG / imp 55 / dev 80
- Uncertainty of Vision Medical Foundation ModelsarXiv cs.LG / imp 45 / dev 65
- Foundation Models Meet Agriculture: Challenges Beyond PretrainingarXiv cs.LG / imp 40 / dev 60
- TopGQ: Fast GNN Post-Training Quantization Leveraging Topology InformationarXiv cs.LG / imp 45 / dev 75
- Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware CriticarXiv cs.LG / imp 40 / dev 70
- No Equivariant Architecture Covers All Equivariant AttentionarXiv cs.LG / imp 40 / dev 65
- Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient RegistryarXiv cs.LG / imp 50 / dev 60
- PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of ExpertarXiv cs.LG / imp 40 / dev 70
- Self-Supervised Pretext Tasks for Infant Cry Analysis: A Controlled Comparison and a Cautionary Result on DonateacryarXiv cs.LG / imp 30 / dev 55
- Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal GeometryarXiv cs.LG / imp 50 / dev 70
- ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity PredictionarXiv cs.LG / imp 40 / dev 65
- Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 FrameworkarXiv cs.LG / imp 45 / dev 65
- When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast StreamsarXiv cs.LG / imp 40 / dev 60
- Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and InterpretabilityarXiv cs.LG / imp 60 / dev 80
- Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport SystemsarXiv cs.LG / imp 30 / dev 60
- PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMsarXiv cs.LG / imp 60 / dev 80
- Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMsarXiv cs.LG / imp 55 / dev 85
- Collapsibility of Performance Metrics in Clinical Predictive AIarXiv cs.LG / imp 40 / dev 65
- The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and RefusalarXiv cs.LG / imp 60 / dev 75
- State of Health Estimation using Convolutional and Bidirectional LSTM Neural Networks tuned by Bayesian OptimizationarXiv cs.LG / imp 25 / dev 55
- PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference OptimizationarXiv cs.LG / imp 55 / dev 80
- MolLedger: An Additive Graph Neural Network with Chemically Grounded ADME AttributionsarXiv cs.LG / imp 35 / dev 65
- Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RLarXiv cs.LG / imp 45 / dev 70
- Season-Aware Hybrid Convolutional-Transformer for Antarctic Sea Ice Concentration ForecastingarXiv cs.LG / imp 30 / dev 60
- CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model SynergyarXiv cs.LG / imp 50 / dev 75
- Learning Materials Properties from Scarce Labels and Unlabeled CrystalsarXiv cs.LG / imp 35 / dev 65
- Liquid Gated AttentionarXiv cs.LG / imp 45 / dev 75
- Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised LearningarXiv cs.LG / imp 40 / dev 70
- Kolmogorov--Arnold against bounded translationsarXiv cs.LG / imp 45 / dev 65
- Tracing distinguishability through transformer processing with stochastic LayerNormarXiv cs.LG / imp 45 / dev 70
- BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML TasksarXiv cs.LG / imp 70 / dev 85
- E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business OperationarXiv cs.LG / imp 75 / dev 85
- Functional Degeneracy in Neural Networks: Measurement and PruningarXiv cs.LG / imp 50 / dev 75
- TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic ClassificationarXiv cs.LG / imp 35 / dev 65
- Do VLMs Share Safety Neurons Across Modalities?arXiv cs.LG / imp 65 / dev 80
- PRACTICE: From Experience to Expertise in Self-Evolving Embodied AgentsarXiv cs.LG / imp 75 / dev 85
- T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and SchedulerarXiv cs.LG / imp 50 / dev 80
- TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model TrainingarXiv cs.LG / imp 65 / dev 85
- Reciprocity Separates Gradient Flow from Rotation in Conservative Physical LearningarXiv cs.LG / imp 40 / dev 60
- Geometric Attractor Monitoring: A Robust and Frugal Framework for Multi-modal Industrial Robotic CyclesarXiv cs.LG / imp 35 / dev 65
- What Emerges and What Breaks in Self-Play DrivingarXiv cs.LG / imp 55 / dev 75
- Deploying DeepSeek 175B Locally on a Single Consumer-Grade RTX 4060 Laptop with 32GB RAM for 200k-Scale Protein-Ligand Virtual ScreeningarXiv cs.LG / imp 65 / dev 85
- Fine-Tuning Low-Bit Models with Gradient in Quantized Code SpacearXiv cs.LG / imp 55 / dev 85
- S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure ElucidationarXiv cs.LG / imp 70 / dev 85
- Selection-Aware Stress Testing for Interactive AgentsarXiv cs.LG / imp 65 / dev 80
- Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning MethodsarXiv cs.LG / imp 40 / dev 70
- Nonparametric Contextual Pricing and Inventory Learning under Censored DemandarXiv cs.LG / imp 35 / dev 60
- Reproducible macroscopic dynamics in a closed-loop human-AI learning systemarXiv cs.LG / imp 60 / dev 75
- One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool LearningarXiv cs.LG / imp 60 / dev 80
- Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not CommutearXiv cs.LG / imp 40 / dev 65
- A Universal Context-Reuse Layer for Cross-Model KV SharingarXiv cs.LG / imp 60 / dev 85
- A Human-in-the-Loop Autonomous Agent for Industry Time Series ForecastingarXiv cs.LG / imp 65 / dev 80
- Sparse Competition during Training For the Emergence of Specialized ModulesarXiv cs.LG / imp 50 / dev 70
- Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation SteeringarXiv cs.LG / imp 65 / dev 80
- Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular GenerationarXiv cs.LG / imp 50 / dev 75
- TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series ClassificationarXiv cs.LG / imp 45 / dev 70
- Normalized Low-Rank AdaptationarXiv cs.LG / imp 55 / dev 85
- Rotational Equivariance in Machine Learning: A Comprehensive TutorialarXiv cs.LG / imp 45 / dev 70
- Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-ImprovementarXiv cs.LG / imp 55 / dev 80
- Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny TransformersarXiv cs.LG / imp 50 / dev 75
- A Model with No Head and Many ThoughtsarXiv cs.LG / imp 65 / dev 85
- Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference OptimizationarXiv cs.LG / imp 60 / dev 80
- Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark ConclusionsarXiv cs.LG / imp 65 / dev 80
- On the Complexity of the Compatibility Problem for Succinctly Encoded Conditional DistributionsarXiv cs.LG / imp 30 / dev 55
- Sharp Approximation Rates for Neural Networks with Affine Latent ParameterizationsarXiv cs.LG / imp 40 / dev 65
- Constant Individual Regret in General GamesarXiv cs.LG / imp 35 / dev 60
- Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFTarXiv cs.LG / imp 70 / dev 90
- Privacy-Preserving Detection of Rare Disease-Associated Cell Subsets via Secure Multi-Party ComputationarXiv cs.LG / imp 45 / dev 65
- Propensity Straight-Through Gradients for Discrete Stochastic SystemsarXiv cs.LG / imp 40 / dev 75
- The Signal in the Noise: An Auditable Reliability Layer for Biomedical Text ClassificationarXiv cs.LG / imp 55 / dev 80
- Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human ValuesarXiv cs.LG / imp 65 / dev 80
- Machine Learning-Enhanced Tabu Search for Tactical Wireless Network DesignarXiv cs.LG / imp 40 / dev 70
- AutoScientist-Quant: Self-Evolving Coding Agents for Automatic Research in Quantitative InvestmentarXiv cs.LG / imp 75 / dev 90
- Reward-Oracle MCTS for Formal Theorem Proving: Sample-Efficient Search and the Need for Kernel-Level Proof AuditingarXiv cs.LG / imp 65 / dev 85
- From Extraction to Governed Memory: Multi-Agent Knowledge Graph Construction with Domain-Expert ReviewarXiv cs.LG / imp 75 / dev 85
- How Language Models Choose Sides: Internal Representations of Instruction HierarchyarXiv cs.LG / imp 65 / dev 80
- Test-Time Scaling for Scientific Equation DiscoveryarXiv cs.LG / imp 65 / dev 85
- FrameScope: Temporal Data Valuation for Stream Active Learning in Autonomous Vehicle SystemsarXiv cs.LG / imp 55 / dev 80
- AdaptAV: Continuous Adaption of Vision Models for Autonomous Vehicles Using Cloud-based OraclearXiv cs.LG / imp 55 / dev 80
- Distributed Semantic Segmentation With Improved Rate-Distortion Trade-OffarXiv cs.LG / imp 45 / dev 75
- Data Diversity, Not Frequency Invariance: A Controlled and Self-Audited Study of Compression-Robust Deepfake DetectionarXiv cs.LG / imp 55 / dev 80
- ORDDAR: Observation-Driven Reasoning for Distortion-Resilient Decision, Action, and Cognitive RecoveryarXiv cs.LG / imp 70 / dev 85
- Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive DiffusionarXiv cs.LG / imp 50 / dev 75
- Adversarial Calibration Attack on Autonomous VehiclesarXiv cs.LG / imp 10 / dev 20
- ASTRA - Agentic System for Ticket Resolution and AnalysisarXiv cs.LG / imp 45 / dev 65
- Separable Nonnegative Matrix Factorization Using Powered Ratio-of-Norms RegularizationarXiv cs.LG / imp 15 / dev 40
- Quantitative Target Convergence and Uniform-in-Time Propagation of Chaos for Langevin-Regularized SVGDarXiv cs.LG / imp 5 / dev 15
- Representation Learning with Quantum Signal ProcessingarXiv cs.LG / imp 10 / dev 30
- Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability AnalysisarXiv cs.LG / imp 20 / dev 50
- A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternativesarXiv cs.LG / imp 50 / dev 75
- Uncertainty-Aware Multi-Task Learning for Joint Modulation Recognition and SINR EstimationarXiv cs.LG / imp 15 / dev 40
- Generative Translation Priors: Bayesian Imaging with Cross-Modality Image TranslationarXiv cs.LG / imp 20 / dev 50
- MWIR-4-Plastic: The Identification of Complex End-of-Life Industrial Plastic using Mid-wave Infrared Hyperspectral Imaging and Machine LearningarXiv cs.LG / imp 15 / dev 40
- Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended GenerationarXiv cs.LG / imp 35 / dev 60
- mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC SynthesisarXiv cs.LG / imp 20 / dev 45
- Leveraging Turn-taking Dynamics for Intent Recognition in Multi-party ConversationsarXiv cs.LG / imp 25 / dev 55
- The Hallucination Signal Is a Mean Shift: Why Simple Probes SufficearXiv cs.LG / imp 45 / dev 75
- MERIT: Mitigating Exposure Bias in Generative XMC for User-Interest Propensity ModelingarXiv cs.LG / imp 35 / dev 65
- Oculi: A Conversational Agentic Platform for Automated Credit Risk AnalysisarXiv cs.LG / imp 50 / dev 70
- The information geometry of product-reference discrete diffusion: Interaction growth complexity and optimal schedulingarXiv cs.LG / imp 10 / dev 25
- From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People SearcharXiv cs.LG / imp 25 / dev 55
- Brain-Language-Action (BLA) Models: Language-Conditioned EEG for Robotics ControlarXiv cs.LG / imp 30 / dev 65
- Efficient GPU Retrieval for Semantic SearcharXiv cs.LG / imp 40 / dev 75
- The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model EraarXiv cs.LG / imp 55 / dev 75
- Jigsaw-CRL: Recovering Global Latent Causal Order from Fragmented Multi-Client InterventionsarXiv cs.LG / imp 15 / dev 35
- Sharp Restricted Isometry Thresholds for Global Minima of Rank-Restricted Matrix LASSOarXiv cs.LG / imp 5 / dev 15
- Spectral-Embedded Operator Learning for Three-Phase Interfacial Flow: A Ternary Cahn-Hilliard-Navier-Stokes BenchmarkarXiv cs.LG / imp 15 / dev 50
- Optimally Selecting Representative Agents from a Metric SpacearXiv cs.LG / imp 12 / dev 30
- Clustering as Approximation by Constrained Projectors: Theory and GuaranteesarXiv cs.LG / imp 20 / dev 45
- Emergent Misalignment Is Not MagicalarXiv cs.LG / imp 50 / dev 75
- APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API WorkflowsarXiv cs.LG / imp 55 / dev 80
- Uniform Statistical Convergence of Empirical Sinkhorn Potentials with Exponential and Polynomial Dependence on the Regularization ParameterarXiv cs.LG / imp 8 / dev 20
- Subtraction-Based Tumor Segmentation and Lesion-Centered pCR Prediction for the MAMA-MIA ChallengearXiv cs.LG / imp 18 / dev 50
- Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph ModelingarXiv cs.LG / imp 25 / dev 60
- AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action ModelsarXiv cs.LG / imp 45 / dev 75
- Validating FKG.in: Soundness Assessment in LLM-Augmented Indian Food KnowledgearXiv cs.LG / imp 30 / dev 60
- QCell: Recombining and Aligning Cell Queries for Overlapping Instance SegmentationarXiv cs.LG / imp 20 / dev 55
- A-MADiff: Attention-Guided Multi-Agent DRL with Diffusion Policies for Memory-Aware Task Orchestration in Mobile AIGC NetworksarXiv cs.LG / imp 30 / dev 70
- Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC NetworksarXiv cs.LG / imp 20 / dev 60
- Signed random Fourier features for fast density estimation with indefinite kernelsarXiv cs.LG / imp 15 / dev 50
- Learning Simple Test-Time Environments for LLM Web AgentsarXiv cs.LG / imp 45 / dev 75
- APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured MeshesarXiv cs.LG / imp 18 / dev 55
- Spectral Analysis for Sparse Matrix Computation: Insights and PotentialarXiv cs.LG / imp 20 / dev 65
- Evaluating Tiny Recursive Models Across Training for Code GenerationarXiv cs.LG / imp 35 / dev 75
- FiLM-GPNet: Geometry-Aware Pseudo-Supervised Phase Restoration with Zero-Shot Generalization for Large Temporal InSAR StacksarXiv cs.LG / imp 18 / dev 55
- Explanations, Prompts, and Formalizations: Arguments for New Norms in LLM-Enabled Mathematical ResearcharXiv cs.LG / imp 40 / dev 60
- A Visual Question Answering Model to Automate Nondestructive Evaluation Image AnalysisarXiv cs.LG / imp 25 / dev 60
- Polis: 3D Self-Supervision at City ScalearXiv cs.LG / imp 25 / dev 65
- Content Exploration Beyond the Feed: Creator Supply and the Shared CorpusarXiv cs.LG / imp 30 / dev 55
- Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human SurrogatesarXiv cs.LG / imp 35 / dev 65
- Benchmark Contamination: A Taxonomy Organized by Defeated MitigationarXiv cs.LG / imp 40 / dev 70
- Deciding When to Decide: Testing Operational Suboptimality Under Distributional ShiftarXiv cs.LG / imp 20 / dev 50
- ARMOR: Manifold-Oriented Training for Adversarially Robust Aerial Object Detection under Data ScarcityarXiv cs.LG / imp 25 / dev 65
- LoGo: Token-Level Dynamic Local-Global AttentionarXiv cs.LG / imp 40 / dev 75
- TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix OptimizationarXiv cs.LG / imp 35 / dev 70
- Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and SecurityarXiv cs.LG / imp 60 / dev 85
- $\mathcal{N}_0$-Foundation: Towards the Age of Tactile IntelligencearXiv cs.LG / imp 35 / dev 70
- Cross-lingual Functional Vectors for Emotion Detection in Large Language ModelsarXiv cs.LG / imp 30 / dev 65
- Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOpsarXiv cs.LG / imp 50 / dev 85
- ButterMamba: Butterworth-Enhanced Spatial-Temporal Mamba for Efficient Traffic Flow PredictionarXiv cs.LG / imp 20 / dev 65
- Transformer-Based Flow Shop Scheduling Using MILP-Generated Training DataarXiv cs.LG / imp 25 / dev 70
- Neural ODE enhanced linear mixed effect models for estimating complex association patterns of time-varying covariates with the marker trajectoryarXiv cs.LG / imp 18 / dev 55
- A Unified Perspective on Conformal Prediction and Wasserstein Distributionally Robust Optimization for Uncertainty QuantificationarXiv cs.LG / imp 25 / dev 55
- The Price of Intelligence: A Quality-Adjusted Price Index for AI ServicesarXiv cs.LG / imp 45 / dev 60
- Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy DistillationarXiv cs.LG / imp 35 / dev 75
- On the Instance Hardness as a Decision Criterion in TinyML SystemsarXiv cs.LG / imp 20 / dev 65
- Continual Test-Time Adaptation via Entropy Sensitivity-Guidance in Strict Online SettingarXiv cs.LG / imp 25 / dev 70
- Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic SegmentationarXiv cs.LG / imp 30 / dev 75
- Hallucination Mitigation for Large Vision-Language Models via Implicit Feature StabilizationarXiv cs.LG / imp 45 / dev 75
- Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer EvidencearXiv cs.LG / imp 40 / dev 75
- When Safety Speaks a Language: A Mechanistic Analysis of Safety-Language Identity Entanglement in LLMsarXiv cs.LG / imp 40 / dev 75
- Data-Driven Design Optimization of Streaming-Potential-Mediated Electrokinetic Transport of Viscoelastic Fluids in MicrochannelsarXiv cs.LG / imp 15 / dev 50
- EDGE: Engine for Deterministic Graph Evaluation through Conversation Simulation from Graph Structured DSL ConfigurationarXiv cs.LG / imp 55 / dev 80
- An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud DetectionarXiv cs.LG / imp 25 / dev 70
- Evolutionary Soups: Evolving Mixture-of-Experts for Multi-Objective LLM AlignmentarXiv cs.LG / imp 45 / dev 75
- Partition-Aware Unlearning for Removing Spurious Correlations in Large Vision-Language ModelsarXiv cs.LG / imp 35 / dev 75
- TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language ModelsarXiv cs.LG / imp 35 / dev 75
- Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning CurriculaarXiv cs.LG / imp 40 / dev 75
- A Deep Latent Variable Framework for Jointly Modeling Missingness, Measurement Error, and HeterogeneityarXiv cs.LG / imp 18 / dev 55
- Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned EvaluationarXiv cs.LG / imp 35 / dev 70
- Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the GuidearXiv cs.LG / imp 40 / dev 75
- Learning Representations through Token Prediction: Geometry, Approximation, and Downstream GuaranteesarXiv cs.LG / imp 30 / dev 70
- When Does a Classifier Help an LLM? Classifier-Guided Prompting and Hybrid Classifier-LLM Models for Credit-Default PredictionarXiv cs.LG / imp 30 / dev 70
- A Hybrid State-Space Approach for Census-Tract Population EstimationarXiv cs.LG / imp 20 / dev 65
- A Simple Transformer Pipeline for Full-Key Side-Channel Attacks on Uncropped DatasetsarXiv cs.LG / imp 25 / dev 70
- Aligning Multi-Trajectory Supervision with Policy Optimization for VLA DrivingarXiv cs.LG / imp 30 / dev 75
- VIBE: Video Instruction-aligned Background music gEnerationarXiv cs.LG / imp 25 / dev 65
- Balancing Privacy, Utility, and Safety in LLM Alignment through Preference OptimizationarXiv cs.LG / imp 45 / dev 75
- The PUR-1 Cyber-Physical Digital TwinarXiv cs.LG / imp 20 / dev 60
- Fairness in multi-class multi-group classification problems via contextial coherent risk measuresarXiv cs.LG / imp 25 / dev 65
- Motus2: A Self-Evolving General World Model for Dexterous ManipulationarXiv cs.LG / imp 45 / dev 80
- A Borel Concept Class of VC Dimension One with a Non-PAC Consistent Learner in ZFCarXiv cs.LG / imp 8 / dev 20
- Using Prosody to Predict Syntactic StructurearXiv cs.LG / imp 20 / dev 55
- Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training SamplearXiv cs.LG / imp 12 / dev 30
- Dec-BFTRL: Squre-Root Regret for Decentralized Online Upper-Linearizable Optimization under Separation Access with Application to Continuous Submodular MaximizationarXiv cs.LG / imp 10 / dev 25
- Strengthening Recursive Constructions for Zero-Error Shannon CapacityarXiv cs.LG / imp 8 / dev 15
- Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation SteeringarXiv cs.LG / imp 40 / dev 75
- Compact and Infinite-Order Error Analysis for Null-Space SVD EstimationarXiv cs.LG / imp 10 / dev 25
- Kathleen Remembers: Length-Invariant One-Shot Recall Without AttentionarXiv cs.LG / imp 30 / dev 70
- Benchmarking External Generalization of SPD Matrix Learning for Resting-State fMRI Connectome PredictionarXiv cs.LG / imp 20 / dev 65
- ObjectSplat: Improving Mesh Fidelity and Interactivity for 3D Scenes via Object-Level Mesh SplattingarXiv cs.LG / imp 25 / dev 70
- Ceiling-Clipped Acceptance Histograms Indicate Stranded Speed-up in Block-Diffusion Speculative DecodingarXiv cs.LG / imp 35 / dev 75
- Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social InteractionsarXiv cs.LG / imp 40 / dev 75
- Generalization as a robust performance property of learning-enabled dynamical systemsarXiv cs.LG / imp 40 / dev 60
- Event-Driven Language Models with Sparse Neural Activity for Neuromorphic HardwarearXiv cs.LG / imp 45 / dev 65
- End-to-End Neural Shrinkage of Indefinite Pairwise Correlation Matrices for Small-Cap-Inclusive PortfoliosarXiv cs.LG / imp 10 / dev 20
- VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMsarXiv cs.LG / imp 55 / dev 65
- Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard CorpusarXiv cs.LG / imp 15 / dev 40
- TSExplorer: An interactive data annotation and exploration tool for time-series dataarXiv cs.LG / imp 30 / dev 50
- Beamforming Design Via GNN in mmWave Cell-Free Massive MIMO Using Sub-6 GHz CSIarXiv cs.LG / imp 20 / dev 50
- Minerals in the Wild: A Hyperspectral-XRF Dataset for Elemental Composition EstimationarXiv cs.LG / imp 25 / dev 45
- Informative Label Missingness in Multiclass Classification Information Geometry and Excess RiskarXiv cs.LG / imp 35 / dev 55
- Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an OraclearXiv cs.LG / imp 50 / dev 65
- Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-TrainingarXiv cs.LG / imp 35 / dev 60
- GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM ReasoningarXiv cs.LG / imp 55 / dev 75
- Quantum-Grassmann-Plucker Token Mixing for Deep Learning-Based Post-Disaster Damage AssessmentarXiv cs.LG / imp 25 / dev 50
- BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMsarXiv cs.LG / imp 60 / dev 70
- What It Costs to Compose, Rebuild, and Correct Precomputed MemoryarXiv cs.LG / imp 55 / dev 75
- Fine-Grained Multi Image Object Hallucination BenchmarkarXiv cs.LG / imp 50 / dev 70
- An Agentic Retrobiosynthesis Framework with Learned Frontier SelectionarXiv cs.LG / imp 50 / dev 65
- SingProbe Technical ReportarXiv cs.LG / imp 60 / dev 75
- Conjoint Audio-to-Spikes Encoding and Processing for Efficient Neuromorphic Speech RecognitionarXiv cs.LG / imp 30 / dev 60
- Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model ContributionsarXiv cs.LG / imp 40 / dev 60
- TopoCompress: Long Context Compression via Graph-Wired Semantic TrajectoriesarXiv cs.LG / imp 65 / dev 75
- Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition SystemsarXiv cs.LG / imp 40 / dev 60
- Safety Screening for Voltage Control in Active Distribution Grids via Distributionally Robust Conformal ScreeningarXiv cs.LG / imp 25 / dev 50
- CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music RepresentationsarXiv cs.LG / imp 35 / dev 60
- Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-TuningarXiv cs.LG / imp 60 / dev 75
- Learning the Geometry of Admissible Hypotheses through Inductive Bias in Training DistributionsarXiv cs.LG / imp 35 / dev 55
- Driving on MemoryarXiv cs.LG / imp 45 / dev 65
- Segmentation of Bovid Dentition Under Imperfect Annotations: A Comparative Study of Convolutional and Attention ModelsarXiv cs.LG / imp 20 / dev 50
- Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research AgentsarXiv cs.LG / imp 70 / dev 75
- Minimax bounds for watermarked and masked recursive discrete distribution estimationarXiv cs.LG / imp 20 / dev 40
- One Adapter, Many Tasks: Task-Conditioned Feature Transformations for Continual LearningarXiv cs.LG / imp 45 / dev 65
- LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware EngineeringarXiv cs.LG / imp 70 / dev 80
- "Train classical, deploy quantum" requires rethinking generalizationarXiv cs.LG / imp 40 / dev 60
- Implementing neural network mixed-effects models in Template Model Builder (TMB)arXiv cs.LG / imp 30 / dev 55
- GFlowNets and variational inferencearXiv cs.LG / imp 45 / dev 65
- Delta-AI: Local objectives for amortized inference in sparse graphical modelsarXiv cs.LG / imp 35 / dev 60
- Expected flow networks in stochastic environments and two-player zero-sum gamesarXiv cs.LG / imp 40 / dev 65
- Amortizing intractable inference in large language modelsarXiv cs.LG / imp 60 / dev 80
- SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning SystemsarXiv cs.LG / imp 50 / dev 75
- Branch Scaling Manifests as Implicit Architectural Regularization for Improving Generalization in Overparameterized ResNetsarXiv cs.LG / imp 40 / dev 60
- Understanding Deep Learning via Notions of RankarXiv cs.LG / imp 45 / dev 65
- Adaptive teachers for amortized samplersarXiv cs.LG / imp 40 / dev 65
- MEGA: Message Passing Neural Networks for Multigraphs with EdGe AttributesarXiv cs.LG / imp 40 / dev 65
- H-FedSN: Personalized Sparse Networks for Efficient and Accurate Hierarchical Federated Learning for IoT ApplicationsarXiv cs.LG / imp 40 / dev 70
- Alert: Learning Trigger Functions for Early Classification of Time Series using Deep-RLarXiv cs.LG / imp 45 / dev 70
- You Do Not Fully Utilize Transformer's Representation CapacityarXiv cs.LG / imp 55 / dev 75
- RSPO: Regularized Self-Play Alignment of Large Language ModelsarXiv cs.LG / imp 70 / dev 80
- Temporal Analysis of NetFlow Datasets for Network Intrusion Detection SystemsarXiv cs.LG / imp 35 / dev 60
- Semantics at an Angle: When Cosine Similarity Works Until It Doesn'tarXiv cs.LG / imp 50 / dev 75
- Graph Representational Learning: When Does More Expressivity Hurt Generalization?arXiv cs.LG / imp 45 / dev 70
- Large-Scale Bayesian Tensor Reconstruction via Approximate Message PassingarXiv cs.LG / imp 35 / dev 60
- Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM FinetuningarXiv cs.LG / imp 70 / dev 80
- Kronecker Factorization Improves Efficiency and Interpretability of Sparse AutoencodersarXiv cs.LG / imp 55 / dev 75
- EquiReg: Equivariance Regularized Diffusion for Inverse ProblemsarXiv cs.LG / imp 45 / dev 70
- Federated Learning for MRI-based BrainAGE: a multicenter study on post-stroke functional outcome predictionarXiv cs.LG / imp 45 / dev 70
- Discrete Compositional Generation via General Soft Operators and Robust Reinforcement LearningarXiv cs.LG / imp 50 / dev 75
- Training-free LLM Verification via Recycling Few-shot ExamplesarXiv cs.LG / imp 60 / dev 75
- A foundation model with multi-variate parallel attention to generate neuronal activityarXiv cs.LG / imp 45 / dev 70
- A Cycle-Consistency Constrained Framework for Dynamic Solution Space Reduction in Noninjective RegressionarXiv cs.LG / imp 35 / dev 60
- A Conditional GAN for Tabular Data Generation with Probabilistic Sampling of Latent SubspacesarXiv cs.LG / imp 50 / dev 70
- StructSynth: Dependency Graphs as Generation Plans for Low-Data Tabular Synthesis with Language ModelsarXiv cs.LG / imp 55 / dev 75
- Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning EnvironmentarXiv cs.LG / imp 70 / dev 85
- Integrating attention into explanation frameworks for language and vision transformersarXiv cs.LG / imp 50 / dev 75
- EEGDM: Learning EEG Representation with Latent Diffusion ModelarXiv cs.LG / imp 40 / dev 70
- SHAKE-GNN: Scalable Hierarchical Kirchhoff-Forest Graph Neural NetworkarXiv cs.LG / imp 45 / dev 75
- Data-to-Energy Stochastic DynamicsarXiv cs.LG / imp 40 / dev 65
- Multi-Marginal Flow Matching with Adversarially Learnt InterpolantsarXiv cs.LG / imp 45 / dev 70
- MolGA: Molecular Graph Adaptation with Pre-trained 2D Graph EncoderarXiv cs.LG / imp 45 / dev 70
- Personalized Treatment Outcome Prediction from Scarce Data via Dual-Channel Knowledge Distillation and Adaptive FusionarXiv cs.LG / imp 45 / dev 65
- Structure-Preserving Physics-Informed Neural Network for the Korteweg--de Vries (KdV) EquationarXiv cs.LG / imp 40 / dev 70
- Deep Reinforcement Learning for Dynamic Origin-Destination Matrix Estimation in Microscopic Traffic Simulations Considering Credit AssignmentarXiv cs.LG / imp 40 / dev 70
- Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-ExpertsarXiv cs.LG / imp 55 / dev 80
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningarXiv cs.LG / imp 55 / dev 75
- SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial VerificationarXiv cs.LG / imp 65 / dev 85
- ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground TrutharXiv cs.LG / imp 70 / dev 85
- Better World Models Can Lead to Better Post-Training PerformancearXiv cs.LG / imp 70 / dev 80
- Kascade: A Practical Sparse Attention Method for Long-Context LLM InferencearXiv cs.LG / imp 70 / dev 85
- KV Admission: Learning What to Write for Efficient Long-Context LLM InferencearXiv cs.LG / imp 70 / dev 85
- Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMsarXiv cs.LG / imp 60 / dev 75
- Federated Personalization of Early-Exit NetworksarXiv cs.LG / imp 50 / dev 75
- Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language ModelsarXiv cs.LG / imp 65 / dev 80
- Securing Time Integrity in Energy IoT Against Clock Drift and Y2K38 FailuresarXiv cs.LG / imp 35 / dev 65
- Semi-supervised CAPP Transformer Learning via Pseudo-labelingarXiv cs.LG / imp 35 / dev 65
- Universal Redundancies in Time Series Foundation ModelsarXiv cs.LG / imp 55 / dev 80
- Least but not Last: Fine-tuning Intermediate Principal Components for Better Performance-Forgetting Trade-OffsarXiv cs.LG / imp 60 / dev 80
- Constrained Group Relative Policy OptimizationarXiv cs.LG / imp 65 / dev 85
- Zero-shot Generalizable Graph Anomaly Detection with Mixture of Riemannian ExpertsarXiv cs.LG / imp 50 / dev 75
- Reverse N-Wise Output-Oriented Testing for AI/ML and Quantum Computing SystemsarXiv cs.LG / imp 60 / dev 80
- Efficient Real-Time Adaptation of ROMs for Unsteady Flows Using Data AssimilationarXiv cs.LG / imp 40 / dev 70
- MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language ModelsarXiv cs.LG / imp 70 / dev 80
- Personalized Group Relative Policy Optimization for Heterogenous Preference AlignmentarXiv cs.LG / imp 65 / dev 85
- Group Resonance Network: Learnable Prototypes and Multi-Subject Resonance for EEG Emotion RecognitionarXiv cs.LG / imp 40 / dev 70
- Grammar of the Wave: Towards Explainable Multivariate Time Series Event Detection via Neuro-Symbolic VLM AgentsarXiv cs.LG / imp 70 / dev 85
- Generalization and Memorization in Rectified FlowarXiv cs.LG / imp 55 / dev 75
- Mixture-Greedy for Online Generative Model Selection: Is UCB Necessary in Diversity-Aware Multi-Armed Bandits?arXiv cs.LG / imp 45 / dev 70
- Identification of Bivariate Causal Directionality Based on Anticipated Asymmetric GeometriesarXiv cs.LG / imp 35 / dev 60
- LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA ModelsarXiv cs.LG / imp 60 / dev 80
- EvoLen: Evolution-Guided Tokenization for DNA Language ModelarXiv cs.LG / imp 55 / dev 80
- Automated Batch Distillation Process Simulation for a Large Hybrid Dataset for Deep Anomaly DetectionarXiv cs.LG / imp 45 / dev 70
- MOONSHOT : A Framework for Multi-Objective Pruning of Vision and Large Language ModelsarXiv cs.LG / imp 70 / dev 85
- REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy AnnotationsarXiv cs.LG / imp 50 / dev 75
- Perturbation Sensitivity of Maximum-Likelihood Pairwise Ranking in Computational Decision SystemsarXiv cs.LG / imp 20 / dev 30
- FG$^2$-GDN: Enhancing Long-Context Gated Delta Networks with Doubly Fine-Grained ControlarXiv cs.LG / imp 55 / dev 85
- Inverting Foundation Models of Brain Function with Simulation-Based InferencearXiv cs.LG / imp 25 / dev 35
- AutoREC: A reinforcement learning platform for equivalent circuit model generationarXiv cs.LG / imp 35 / dev 60
- AirFM-DDA: Air-Interface Foundation Model in the Delay-Doppler-Angle Domain for AI-Native 6GarXiv cs.LG / imp 40 / dev 70
- Concepts Whisper: Spectral Anti-Concentration and the Dual Geometry of Transformer RepresentationsarXiv cs.LG / imp 55 / dev 80
- Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token LevelarXiv cs.LG / imp 60 / dev 75
- PairAlign: A Framework for Autoregressive Tokenization via Self-Alignment with Applications to Audio TokenizationarXiv cs.LG / imp 55 / dev 75
- OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio ScalingarXiv cs.LG / imp 40 / dev 70
- ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language ModelsarXiv cs.LG / imp 55 / dev 80
- Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought ReasoningarXiv cs.LG / imp 60 / dev 75
- Emulating the Forced Response of Climate Models with Generative Machine LearningarXiv cs.LG / imp 35 / dev 40
- World Model Control by Trajectory Reachability MetricsarXiv cs.LG / imp 45 / dev 70
- Generalist Graph Anomaly Detection via Prototype-Based DistillationarXiv cs.LG / imp 40 / dev 65
- When the Strongest Teacher Is Not the Best Teacher: Student-Centric Answer SelectionarXiv cs.LG / imp 60 / dev 75
- Locality-Aware Redundancy Pruning for LLM Depth CompressionarXiv cs.LG / imp 60 / dev 80
- SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic DenoisingarXiv cs.LG / imp 45 / dev 60
- OISD: On-Policy Internal Self-Distillation of Language ModelsarXiv cs.LG / imp 60 / dev 75
- LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-TrainingarXiv cs.LG / imp 65 / dev 80
- Riemannian Optimization for Hadamard Products of Low-Rank MatricesarXiv cs.LG / imp 30 / dev 50
- Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference MisalignmentarXiv cs.LG / imp 65 / dev 70
- GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-TuningarXiv cs.LG / imp 50 / dev 75
- Mamba-Assisted Non-Markovian Closure for Reduced-Order ModelingarXiv cs.LG / imp 35 / dev 60
- GRASP: Geometry-aware Residual Alignment for Scalable Pretraining Data AttributionarXiv cs.LG / imp 60 / dev 75
- TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series GenerationarXiv cs.LG / imp 30 / dev 50
- QueryGraph: Reliable Multi-Tool Query Execution Planning via LLM-Based Graph GenerationarXiv cs.LG / imp 65 / dev 80
- Divide-and-Conquer Modeling for the CTF-4-Science Lorenz BenchmarkarXiv cs.LG / imp 25 / dev 45
- When Design Rules Break: Benchmark Composition Determines Whether Label Informativeness Predicts GNN Aggregator ChoicearXiv cs.LG / imp 45 / dev 60
- Bergson: An Open Source Library for Data AttributionarXiv cs.LG / imp 60 / dev 80
- A Zero-shot Generalized Graph Anomaly Detection Framework via Node ReconstructionarXiv cs.LG / imp 45 / dev 65
- CARE: Context-Aware Ranking Evolution with Executable Scoring Programs for Budgeted Reaction OptimizationarXiv cs.LG / imp 55 / dev 75
- Learning Generated Controls under Fractured Geometry: Projective Residualization and Variation-Allocation FrontiersarXiv cs.LG / imp 25 / dev 40
- WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource HardwarearXiv cs.LG / imp 70 / dev 85
- NeuReasoner: Theory-grounded Mapping of Reasoning Elicitation BoundariesarXiv cs.LG / imp 65 / dev 75
- Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty EstimationarXiv cs.LG / imp 60 / dev 80
- A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-TrainingarXiv cs.LG / imp 70 / dev 85
- A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent StrongarXiv cs.LG / imp 50 / dev 70
- Activation Steering Transfer to Agents: One Gain Ratio Does Not Identify Potency and EfficacyarXiv cs.LG / imp 60 / dev 75
- PRISM Edit: One Vector for All Temporal AnswersarXiv cs.LG / imp 60 / dev 75
- Decoupled Structure-Feature Alignment via Alternating Optimization for Graph LearningarXiv cs.LG / imp 45 / dev 65
- Energy-Based Physics-Informed Form Finding for Clustered Tensegrity StructuresarXiv cs.LG / imp 25 / dev 40
- Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular DataarXiv cs.LG / imp 50 / dev 65
- Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset GeneratorarXiv cs.LG / imp 65 / dev 85
- Off-Context GRPO: Learning to Reason on Hard Problems using Privileged InformationarXiv cs.LG / imp 70 / dev 80
- Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family EffectsarXiv cs.LG / imp 60 / dev 80
- Error Certificates for KV-Cache Eviction via Randomized DesignarXiv cs.LG / imp 65 / dev 85
- Learning What Matters: Supervising Global Context Pruning with Causal Evidence SetsarXiv cs.LG / imp 70 / dev 80
- Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice CloningarXiv cs.LG / imp 45 / dev 60
- A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor NetworksarXiv cs.LG / imp 45 / dev 65
- SERUM: State Extraction and Refinement for User ModelingarXiv cs.LG / imp 60 / dev 75
- An Identifiability Theory of Masked Prediction: Mode Blindness and Mask SchedulesarXiv cs.LG / imp 40 / dev 65
- Conformalized Large Language Models under Configuration ShiftarXiv cs.LG / imp 65 / dev 80
- Simulation-free and finite-time diffusion modelarXiv cs.LG / imp 45 / dev 70
- Approximate Speculative DecodingarXiv cs.LG / imp 70 / dev 85
- CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure PredictionarXiv cs.LG / imp 50 / dev 70
- From Uncertainty to Failure Attribution: Self-Diagnosing Models for Failure Attribution under Distribution ShiftarXiv cs.LG / imp 60 / dev 75
- Support Selection Beyond Smooth DAG Exactness: Completion Geometry,Score Margins, and Selective CertificatesarXiv cs.LG / imp 25 / dev 40
- Correlation flow governs learning at criticalityarXiv cs.LG / imp 40 / dev 65
- Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal RoutingarXiv cs.LG / imp 75 / dev 85
- When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space CompositionarXiv cs.LG / imp 60 / dev 75
- Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified ConstructionarXiv cs.LG / imp 30 / dev 45
- Scaling Automatic Research Agents via World ModelsarXiv cs.LG / imp 75 / dev 85
- Momentum as Residual-Driven Multiplier Correction for Deep Learning OptimizationarXiv cs.LG / imp 45 / dev 70
- Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM InferencearXiv cs.LG / imp 70 / dev 85
- Large Discovery Models: Empirically-grounded Model-Based Open-Ended SearcharXiv cs.LG / imp 75 / dev 80
- Machine Learning and ARIMA Model Averaging for Adaptive Public Health Forecasting: Comparative Evaluation and an Ontario COVID-19 Case StudyarXiv cs.LG / imp 35 / dev 50
- In-Cell Learning: Language Models That Update Their Own Weights in Sequence Without Changing the File They ShiparXiv cs.LG / imp 65 / dev 85
- Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private DataarXiv cs.LG / imp 75 / dev 80
- Who Should Teach? Confidence-Aware Dual-Teacher Learning for Few-Shot Node Classification on Text-Attributed GraphsarXiv cs.LG / imp 50 / dev 70
- Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for MoleculesarXiv cs.LG / imp 55 / dev 80
- CatchBench: When Can an Agent Failure Be Caught?arXiv cs.LG / imp 70 / dev 80
- Every Layer Counts: An Exponential $L_2$ Depth Hierarchy for ReLU NetworksarXiv cs.LG / imp 40 / dev 60
- Enhancing Bayesian Optimization and Active Learning Through Kernel DiversityarXiv cs.LG / imp 50 / dev 70
- FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM InferencearXiv cs.LG / imp 65 / dev 85
- Demystifying Reinforcement Learning Post-Training of Language ModelsarXiv cs.LG / imp 75 / dev 85
- DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task LearningarXiv cs.LG / imp 35 / dev 55
- MoPLEx: Estimating Plackett-Luce Mixture Models for Multi-Objective AlignmentarXiv cs.LG / imp 60 / dev 75
- Frequency-aware forecasting for short-term typhoon gust predictionarXiv cs.LG / imp 30 / dev 50
- It's a matter of timescale: non-linear utility in successor features and multi-objective planning and learningarXiv cs.LG / imp 55 / dev 70
- ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level ExplanationsarXiv cs.LG / imp 60 / dev 80
- A Unified Framework for Fair and Personalized Decentralized Learning under Communication ConstraintsarXiv cs.LG / imp 55 / dev 75
- QuantumBoostNet: Hybrid Classical-Quantum Cardiac View IdentificationarXiv cs.LG / imp 30 / dev 50
- Beyond Parallel Blindness: Information Floors and Model Gaps in Block DraftingarXiv cs.LG / imp 65 / dev 80
- Deep graph kernel point processes over networksarXiv cs.LG / imp 40 / dev 65
- Simulation-Based Evaluation of Energy-Constrained Quantum-Classical CompetitionarXiv cs.LG / imp 25 / dev 40
- PhyloGFN: Phylogenetic inference with generative flow networksarXiv cs.LG / imp 35 / dev 60
- Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD LearningarXiv cs.LG / imp 55 / dev 75
- PQMass: Probabilistic Assessment of the Quality of Generative Models using Probability Mass EstimationarXiv cs.LG / imp 50 / dev 70
- Model Selection and Parameter Estimation of One-Dimensional Gaussian Mixture ModelsarXiv cs.LG / imp 35 / dev 55
- Learning diverse attacks on large language models for robust red-teaming and safety tuningarXiv cs.LG / imp 75 / dev 85
- Gender, Race, and Intersectional Bias in Resume Screening via Language Model RetrievalarXiv cs.LG / imp 65 / dev 75
- Autoencoders in Function SpacearXiv cs.LG / imp 45 / dev 65
- Classification Drives Geographic Bias in Street Scene SegmentationarXiv cs.LG / imp 55 / dev 70
- Perforated Backpropagation: A Neuroscience Inspired Extension to Artificial Neural NetworksarXiv cs.LG / imp 40 / dev 60
- An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network PruningarXiv cs.LG / imp 60 / dev 80
- Trust Under Siege: Label Spoofing Attacks against Machine Learning for Android Malware DetectionarXiv cs.LG / imp 65 / dev 75
- Mirror Descent Linearized Augmented Lagrangian Methods for Nonconvex Constrained Stochastic Zeroth-Order OptimizationarXiv cs.LG / imp 35 / dev 55
- mRNA Design and Optimization with Deep Knowledge-Infused ApproacharXiv cs.LG / imp 65 / dev 80
- Optimal Estimation of Watermark Proportions in Hybrid AI-Human TextsarXiv cs.LG / imp 60 / dev 75
- Test of partial effects for Frechet regression on Bures-Wasserstein manifoldsarXiv cs.LG / imp 5 / dev 10
- PERK: Long-Context Reasoning as Test-Time LearningarXiv cs.LG / imp 55 / dev 70
- Language-Guided Tuning: Configuration Optimization for Automated ML ResearcharXiv cs.LG / imp 40 / dev 75
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety InjectionarXiv cs.LG / imp 60 / dev 65
- SPADE: A Large Language Model Framework for Soil Moisture Pattern Recognition and Anomaly Detection in Precision AgriculturearXiv cs.LG / imp 15 / dev 30
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesarXiv cs.LG / imp 35 / dev 85
- TPSO: Training-Free Diverse Image Generation via Semantic Prompt Embedding OptimizationarXiv cs.LG / imp 30 / dev 60
- SpIDER: Spatially Informed Dense Embedding Retrieval for Software Issue LocalizationarXiv cs.LG / imp 65 / dev 85
- Diffusion Models in Simulation-Based Inference: A Tutorial ReviewarXiv cs.LG / imp 40 / dev 65
- Fitted Q-Evaluation without Bellman Completeness via Occupancy WeightingarXiv cs.LG / imp 25 / dev 60
- Soft Fitted Q-Iteration without Bellman Completeness: Occupancy Reweighting and Temperature AnnealingarXiv cs.LG / imp 25 / dev 60
- X-Coder: Advancing Competitive Programming with Synthetic Tasks, Solutions, and TestsarXiv cs.LG / imp 55 / dev 80
- Tracing the Latent Threads: A Mechanistic Study of How LLMs Represent and Operationalize Race and Ethnicity CuesarXiv cs.LG / imp 50 / dev 55
- Social Caption: Evaluating Social Understanding in Multimodal ModelsarXiv cs.LG / imp 35 / dev 60
- Learning to Optimize by Differentiable ProgrammingarXiv cs.LG / imp 50 / dev 80
- Unknown Unknowns: Do Hidden Intentions in LLMs Evade Detection?arXiv cs.LG / imp 60 / dev 50
- Diverse via bounded Agreement: Geometric Regularization for Multimodal FusionarXiv cs.LG / imp 30 / dev 75
- Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent ReasoningarXiv cs.LG / imp 50 / dev 70
- Efficient reduction of stellar contamination and noise in planetary transmission spectra using neural networksarXiv cs.LG / imp 5 / dev 10
- Edge-Local and Qubit-Efficient Quantum Graph Learning for the NISQ EraarXiv cs.LG / imp 20 / dev 70
- AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable ApproacharXiv cs.LG / imp 30 / dev 65
- DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design GenerationarXiv cs.LG / imp 30 / dev 65
- Prediction-Powered Conditional InferencearXiv cs.LG / imp 35 / dev 75
- NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document RetrievalarXiv cs.LG / imp 45 / dev 80
- PA3: Policy-Aware Agent Alignment through Chain-of-ThoughtarXiv cs.LG / imp 65 / dev 70
- Interpretable Predictability-Based AI Text Detection: A Replication StudyarXiv cs.LG / imp 40 / dev 65
- Model Selection and Parameter Estimation for Multidimensional Gaussian Mixture Models with a Common Covariance MatrixarXiv cs.LG / imp 20 / dev 65
- Accelerate Vector Diffusion Maps by LandmarksarXiv cs.LG / imp 25 / dev 70
- PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded VerificationarXiv cs.LG / imp 30 / dev 60
- Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and SpoofingarXiv cs.LG / imp 40 / dev 75
- Operator Learning for Predicting Bulk Wave Parameters of Spectral Wave ModelsarXiv cs.LG / imp 20 / dev 65
- SynMulti: Synthetic-to-Real Learning for Multimodal Video UnderstandingarXiv cs.LG / imp 45 / dev 70
- VeriX-Anon: A Multi-Layered Framework for Mathematically Verifiable Outsourced Target-Driven Data AnonymizationarXiv cs.LG / imp 35 / dev 70
- Performance Manipulation: Labor Market Implications in AI-assisted EraarXiv cs.LG / imp 50 / dev 30
- Large language model-enabled automated data extraction for concrete materials informaticsarXiv cs.LG / imp 40 / dev 70
- When Chain-of-Thought Fails, the Solution Hides in the Hidden StatesarXiv cs.LG / imp 55 / dev 75
- Learning the Channel Gain from Anywhere to Anywhere via Cross-environment Transformer EstimatorsarXiv cs.LG / imp 15 / dev 65
- Robust Multi-Agent LLMs under Byzantine FaultsarXiv cs.LG / imp 60 / dev 75
- Operator-Guided Model Reduction for Generative Sampling in Lattice Field TheoryarXiv cs.LG / imp 20 / dev 70
- Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low LightarXiv cs.LG / imp 30 / dev 70
- SciAtlas: A Computable Atlas of Science for Knowledge-Grounded AI ResearcharXiv cs.LG / imp 55 / dev 75
- Physics-Guided Concentration Inference from Resistance Transients in a Mixed-Phase SnO-SnO$_2$ Carbon Monoxide Sensor with p-n SwitchingarXiv cs.LG / imp 20 / dev 70
- Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language ModelsarXiv cs.LG / imp 45 / dev 65
- Measuring the Depth of LLM Unlearning via Activation PatchingarXiv cs.LG / imp 50 / dev 70
- Correcting test set contamination by spiking the training dataarXiv cs.LG / imp 40 / dev 75
- Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation ExplanationarXiv cs.LG / imp 45 / dev 75
- Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-TrainingarXiv cs.LG / imp 55 / dev 75
- Extracting Small Translation Specialists from LLMs by Aggressively Pruning ExpertsarXiv cs.LG / imp 50 / dev 80
- Privacy-Enhanced Zero-Order Federated Learning via xMK-CKKS over Wireless ChannelsarXiv cs.LG / imp 40 / dev 80
- When Should Models Change Their Minds? Contextual Belief Management in Large Language ModelsarXiv cs.LG / imp 50 / dev 70
- Exploring Autonomous Agentic Data Engineering for Model SpecializationarXiv cs.LG / imp 65 / dev 80
- Self-Correction Can Amplify Hallucinations: Fact-Level Repair with Graph-Based Evidence Routing in Multimodal GenerationarXiv cs.LG / imp 50 / dev 70
- DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion ModelsarXiv cs.LG / imp 35 / dev 80
- Don't Read Everything: A Curvature-Conditioned Query for Linear AttentionarXiv cs.LG / imp 40 / dev 85
- Evolving Agents in the Dark: Retrospective Harness Optimization via Self-PreferencearXiv cs.LG / imp 75 / dev 80
- CrowdMath: A Dataset of Crowdsourced Mathematical Research DiscussionsarXiv cs.LG / imp 40 / dev 60
- Twelve quick tips for designing AI-driven HPC workflowsarXiv cs.LG / imp 55 / dev 80
- How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial TransactionsarXiv cs.LG / imp 50 / dev 85
- From AGI to ASIarXiv cs.LG / imp 70 / dev 40
- PEAR: Permutation-Equivariant Adaptive Routing Multi-Agent DebatearXiv cs.LG / imp 60 / dev 80
- In LLM Reasoning, there is Irrationality on top of Value MisalignmentarXiv cs.LG / imp 55 / dev 70
- FracEvent: Event-Camera Simulation via Fractional-Relaxation Pixel DynamicsarXiv cs.LG / imp 25 / dev 75
- MultiHashFormer: Hash-based Generative Language ModelsarXiv cs.LG / imp 50 / dev 85
- J-LAW: Joint Localization and Action-Conditioned World Modeling via Coupled Latent Factor GraphsarXiv cs.LG / imp 35 / dev 75
- Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group DiscoveryarXiv cs.LG / imp 35 / dev 75
- Cultural Bias Without a Cultural Self:A Disassociation Study of LLM's Persona and BiasarXiv cs.LG / imp 45 / dev 55
- Runtime Safety Filtering for Learned Small UAS Separation Policies under GNSS DegradationarXiv cs.LG / imp 40 / dev 80
- Posterior Variance Is a Constraint Map, Not an Error Map: Closed-Form Uncertainty for Radiative Gaussian Splatting in Sparse-View CTarXiv cs.LG / imp 25 / dev 75
- RegionFM: Interpretable Region-Based Brain MRI Classification Using Foundation Model EmbeddingsarXiv cs.LG / imp 30 / dev 65
- The Value of Depth in Message Passing on Sparse Graphs: A Kesten-Stigum DichotomyarXiv cs.LG / imp 30 / dev 80
- SALT: Salience-Aware Lexical Trie for Long-Context CompressionarXiv cs.LG / imp 60 / dev 80
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football ForecastingarXiv cs.LG / imp 50 / dev 70
- Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World ModelsarXiv cs.LG / imp 40 / dev 75
- Learning to Trace Seiberg DualitiesarXiv cs.LG / imp 20 / dev 65
- Fast Trainable Multilinear Bases for Image CompressionarXiv cs.LG / imp 30 / dev 80
- Dynamically Allocating Evaluation Effort for Model RankingarXiv cs.LG / imp 45 / dev 75
- SDF-Aware Weighting: Adaptive Eikonal Regularisation for Three-Dimensional Level-Set Physics-Informed Neural NetworksarXiv cs.LG / imp 25 / dev 75
- MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering ModelarXiv cs.LG / imp 65 / dev 80
- A Deterministic Constant-Competitive Algorithm for Dynamic Mixture-of-Experts ServingarXiv cs.LG / imp 45 / dev 85
- A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous DesignarXiv cs.LG / imp 45 / dev 70
- DELE-w0.5: Inferring Action from Future Latent State for Robotic ManipulationarXiv cs.LG / imp 40 / dev 80
- GTA-RAG: Graph-Trajectory-Augmented Reinforcement Learning for Multi-Turn Retrieval-Augmented ReasoningarXiv cs.LG / imp 70 / dev 85
- LUCAID: Agentic Multimodal AI for Lung Cancer Precision PathologyarXiv cs.LG / imp 50 / dev 75
- SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed LearningarXiv cs.LG / imp 35 / dev 80
- Forecasting Weather-Driven Price Dynamics Across Sri Lankan Tea Market CataloguesarXiv cs.LG / imp 20 / dev 60
- Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement LearningarXiv cs.LG / imp 55 / dev 80
- Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case StudyarXiv cs.LG / imp 50 / dev 80
- STEP: A Modular Silent Trial Engine for Operational Evaluation of Digital Pathology AI in Routine WorkflowarXiv cs.SE / imp 45 / dev 70
- Rust's Type Checker Implementation Is Unsound: An Empirical Study on Soundness Bugs in rustcarXiv cs.SE / imp 55 / dev 85
- Beyond Vector Search: Comparing Classical RAG with Hybrid GraphRAG for Climate Science Q\&AarXiv cs.SE / imp 65 / dev 85
- The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agentsarXiv cs.SE / imp 70 / dev 85
- UML Class Diagram Evaluation and Repair Strategies based on LLMsarXiv cs.SE / imp 40 / dev 80
- FlowCheck: Helping End-Users Specify and Verify Intent in Vibe-Coded Web AppsarXiv cs.SE / imp 50 / dev 75
- Legacy System Modernization with Coding Agents: A Case StudyarXiv cs.SE / imp 65 / dev 85
- AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud AgentarXiv cs.SE / imp 70 / dev 85
- Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single CodebasearXiv cs.SE / imp 65 / dev 85
- Drive the Thoughts: Runtime Monitoring of VLA Reasoning-Trajectory ConsistencyarXiv cs.SE / imp 50 / dev 80
- InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed InformationarXiv cs.SE / imp 45 / dev 75
- Cost-Effective Repository Exploration for Agentic Issue LocalizationarXiv cs.SE / imp 70 / dev 85
- Agent-Driven Verification of Memory Safety for liblzma Decoder Components with VSTarXiv cs.SE / imp 55 / dev 85
- A Comprehensive Study of Native Code Bugs in Python ApplicationsarXiv cs.SE / imp 45 / dev 75
- Open-Source Autonomous Driving System Analysis and Multi-Disciplinary Hardware-in-the-Loop Research Paradigm with Reinforcement-Learning Testing and Large Language ModelsarXiv cs.SE / imp 40 / dev 60
- DSEffi-Bench: Demystifying Large Language Models' Capability in Efficient Data Science Code GenerationarXiv cs.SE / imp 65 / dev 80
- Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?arXiv cs.SE / imp 75 / dev 85
- Bridge: Automatically Mining Ecosystem-Scale API Update Mappings and Client Update InstancesarXiv cs.SE / imp 60 / dev 80
- Developer Attitudes and Practices Towards Optimizing Software Energy ConsumptionarXiv cs.SE / imp 35 / dev 65
- Practical Implementation Report on Introducing Spec-Driven Development Using AI Agents in Software Development PBLarXiv cs.SE / imp 75 / dev 85
- A Phased Workflow for Operating LLM-Based Coding AgentsarXiv cs.SE / imp 80 / dev 90
- sbom-unifier: Integration Framework for Heterogeneous SBOMsarXiv cs.SE / imp 50 / dev 75
- On the Prospects of Dynamic LLM Conversations in Software DevelopmentarXiv cs.SE / imp 65 / dev 85
- ProofPulse: Interactive Proof Coverage Analysis for DafnyarXiv cs.SE / imp 40 / dev 75
- Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity VerificationarXiv cs.SE / imp 70 / dev 80
- MIRAGE-CAD: Construction-Mediated Multimodal Generation of Executable CAD ProgramsarXiv cs.SE / imp 45 / dev 65
- FoldKit: A Python library for efficient storage and retrieval of co-folding predictionsarXiv cs.SE / imp 35 / dev 70
- Emergent Behavior and Uncertainty in IoT-Enhanced Business Processes: Challenges and Future DirectionsarXiv cs.SE / imp 40 / dev 55
- The Web-CLI: Verifiable Privacy for Tools, Models, and Inference Engines in the BrowserarXiv cs.SE / imp 65 / dev 80
- Towards Fully Automated Medical Imaging Code Generation via Validation-based Context EngineeringarXiv cs.SE / imp 55 / dev 75
- A Multi-Month Study of Git Commit SigningarXiv cs.SE / imp 50 / dev 80
- Database-Augmented RAG for Automated Repair of REST API MisusesarXiv cs.SE / imp 60 / dev 80
- EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI AssistantsarXiv cs.SE / imp 65 / dev 85
- Evaluating a 4B open-weights local LLM for agentic DFT workflows: a literature reproducibility auditarXiv cs.SE / imp 65 / dev 80
- Building the Truman Show: A TrustZone-Based Framework for Lightweight Out-of-band Kernel Security MonitoringarXiv cs.SE / imp 60 / dev 75
- POLYFLOW: A Neuro-Symbolic Framework for Static Cross-Language Information Flow AnalysisarXiv cs.SE / imp 60 / dev 80
- A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization AgentsarXiv cs.SE / imp 70 / dev 85
- Verification-Time Dependency on a Disappearing EvaluatorarXiv cs.SE / imp 70 / dev 65
- ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive AlternativesarXiv cs.SE / imp 65 / dev 75
- Detecting DBMS Bugs by Constructing Equivalent Representations of Intermediate Query ResultsarXiv cs.SE / imp 55 / dev 75
- WebWorld: The Browser as a World Model for Self-Improving Web CodearXiv cs.SE / imp 75 / dev 85
- Designing an Auditable LLM-Supported Workflow for Qualitative Thematic AnalysisarXiv cs.SE / imp 45 / dev 60
- Lie to Me: Finding Bugs in ZK DSL Toolchains with Adversarial Witness InjectionarXiv cs.SE / imp 60 / dev 80
- LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent WorkflowarXiv cs.SE / imp 75 / dev 90
- Schwarz: Solver-Aware Agentic Program VerificationarXiv cs.SE / imp 70 / dev 85
- The Exclusion Ratchet: False-Positive Suppression Accumulates and Persists in Detection Rule RepositoriesarXiv cs.SE / imp 55 / dev 70
- MaCTG: Multi-Agent Collaborative Thought Graph for Automatic ProgrammingarXiv cs.SE / imp 75 / dev 90
- Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical StudyarXiv cs.SE / imp 75 / dev 85
- TRACE: Evaluating Execution Efficiency of LLM-Based Code TranslationarXiv cs.SE / imp 70 / dev 85
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code QualityarXiv cs.SE / imp 75 / dev 90
- LogICL: Distilling LLM Reasoning to Bridge the Semantic Gap in Cross-Domain Log Anomaly DetectionarXiv cs.SE / imp 60 / dev 75
- An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific ProcessarXiv cs.SE / imp 50 / dev 70
- Interactive Clarification for Cloud Infrastructure-as-Code SynthesisarXiv cs.SE / imp 70 / dev 85
- A Longitudinal Study of Dependency Reclassifications in JavaScript ProjectsarXiv cs.SE / imp 50 / dev 75
- ReproBreak: A Dataset of Reproducible Web Locator BreaksarXiv cs.SE / imp 60 / dev 75
- How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World SessionsarXiv cs.SE / imp 85 / dev 90
- DeployBench: Benchmarking LLM Agents for Research Artifact DeploymentarXiv cs.SE / imp 70 / dev 85
- Skills for the future software profession: beyond agentic AI!arXiv cs.SE / imp 75 / dev 85
- Same Scrutiny, More Time: Eye Tracking Insights into Reviewing LLM-Labelled CodearXiv cs.SE / imp 65 / dev 80
- ClarifyCodeBench: Evaluating LLMs on Clarifying Ambiguous Requirements for Code GenerationarXiv cs.SE / imp 70 / dev 85
- Industrial Practice of LLM-Based Test Case Carving and Assertion Generation (Experience Paper)arXiv cs.SE / imp 75 / dev 90
- From C to Idiomatic Rust: A Ship-of-Theseus Agentic TranslationarXiv cs.SE / imp 80 / dev 90
- Self-Evolving Coding AgentsarXiv cs.SE / imp 85 / dev 90
- Improving Debugging in Verification-Aware Languages Through Automated Fault Localization: A Case Study in DafnyarXiv cs.SE / imp 65 / dev 80
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core EvolutionarXiv cs.SE / imp 90 / dev 95
- What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect EvalsarXiv cs.SE / imp 65 / dev 75
- Callability Is Not Operability: Controlled Interface Interventions for LLM AgentsarXiv cs.SE / imp 75 / dev 85
- Characterising Global Platforms: Centralised, Decentralised, Federated, and GrassrootsarXiv cs.SE / imp 40 / dev 55
- Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact VerificationarXiv cs.SE / imp 55 / dev 65
- Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent SystemsarXiv cs.SE / imp 75 / dev 85
- Fine, I’ll build my own text editorLobsters / imp 5 / dev 20
- Zuzai, a new word, indicates the absence of AILobsters / imp 5 / dev 10
- Is Minifying CSS Necessary? (2023)Lobsters / imp 10 / dev 40
- A bicycle for the mindLobsters / imp 5 / dev 10
- Janet 1.42.0Lobsters / imp 25 / dev 60
- A Crash Course in Predicate LogicLobsters / imp 20 / dev 50
- This Month in KDE Linux: August 2026Lobsters / imp 25 / dev 40
- RangeFrom, Part 2..: What I think is wrong about the designLobsters / imp 30 / dev 60
- Breaking down Amazon’s mega dropdown (2013)Lobsters / imp 10 / dev 35
- Wasmi 2.0 - Engineering of the Fastest Wasm InterpretersLobsters / imp 60 / dev 80
- curl: a CVE disputeLobsters / imp 40 / dev 70
- "iT woRKs BeTter in THe aPp!!"Lobsters / imp 5 / dev 15
- VibeCoded AI-Slop License v1.0Lobsters / imp 5 / dev 10
- I attended a conference recently and AI use by academics was absurdLobsters / imp 30 / dev 40
- Coherence and orphan instance rulesLobsters / imp 30 / dev 65
- On not becoming a cyborgLobsters / imp 5 / dev 10
- July in Servo: more platforms, faster canvas, web fonts in SVG, and moreLobsters / imp 40 / dev 65
- The Robot Framework languageLobsters / imp 40 / dev 70
- Texttile, a multiplayer blog engine for people who write togetherLobsters / imp 20 / dev 50
- Let’s Use the Emergent CSS random() Function in all the BrowsersLobsters / imp 25 / dev 55
- Bootstrappable builds: how and whyLobsters / imp 50 / dev 75
- V0.1.0.0 of ghcup-gtk releasedLobsters / imp 15 / dev 50
- A Better SQL in 11 Lines of CodeLobsters / imp 30 / dev 65
- Are We Legacy Computing Yet?Lobsters / imp 35 / dev 55
- Revenue Strategies for AI API ServicesDev.to AI / imp 15 / dev 25
- The AI Attack Surface Your Application Security Checklist Does Not CoverDev.to AI / imp 70 / dev 85
- Biomarker Tracking AI: Essential Continuous InsightsDev.to AI / imp 35 / dev 50
- From Designer to AI Automation Engineer: My Journey into AIDev.to AI / imp 25 / dev 40
- I published my first calculator onlineDev.to AI / imp 10 / dev 30
- Building a Crypto Signal Bot with AI APIs - 2026 GuideDev.to AI / imp 40 / dev 70
- Building Agent X: An Autonomous AI Workspace for Developers & Creators 🚀Dev.to AI / imp 35 / dev 60
- An AI's Completely Ordinary Day (A True Story)Dev.to AI / imp 5 / dev 15
- OpenClaw 2.0 Guide: How to Use It, Best Prompts & Use Cases (2026)Dev.to AI / imp 65 / dev 85
- Constitutional Methods for LLMs: Turning Written Principles into Training SignalsDev.to AI / imp 70 / dev 80
- Rewind or Fork? Two Ways to Recover a SolonCode ConversationDev.to AI / imp 60 / dev 75
- AI Agents - Introduction to LLM and AI TerminologiesDev.to AI / imp 25 / dev 50
- Why I designed my compiler to emit only LLVM IR instead of binaries: Hitchhiking Clang and zero-allocation iteratorsDev.to LLM / imp 55 / dev 80
- Using LLMs for Crypto Market Analysis in 2026Dev.to LLM / imp 45 / dev 70
- Would your RAG eval suite notice if someone weakened the prompt?Dev.to LLM / imp 70 / dev 85
- I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single EditDev.to LLM / imp 80 / dev 90
- The Agent Knew It Was Wrong. The System Let It ShipDev.to LLM / imp 80 / dev 85
- 48-Hour Prompt Injection Audit: Field Notes from a Free Model on a Free ServerDev.to LLM / imp 70 / dev 85
- Human-in-the-Loop for Personal AI AssistantsDev.to LLM / imp 75 / dev 85
- Human Approval for Solo-Founder AI AutomationsDev.to LLM / imp 40 / dev 55
- You're Not Using Enough Guardrails — Here's What Actually Works (1788283427249)Dev.to LLM / imp 45 / dev 70
- Fairness Under the Microscope: Why HY3 Beats Nemotron 3 Ultra on lforla's Bias Stereotypes AuditDev.to LLM / imp 25 / dev 40
- Discussion Hub for new Claude incident: Degraded performance on platform.claude.com and Claude for Microsoft Office 365 on Sep 1, 2026r/ClaudeAI / imp 15 / dev 30
- Presenting my dumbest idea yet. The Claw’deck.r/ClaudeAI / imp 5 / dev 25
- It seems we just got a limit reset with the release of Fable 5.1r/ClaudeAI / imp 20 / dev 35
- “Honey, I upgraded us to the 20x plan, so I get 20x the action now, right? …Right?”r/ClaudeAI / imp 5 / dev 20
- I asked Claude to draw itself after analyzing our chat history.r/ClaudeAI / imp 5 / dev 15
- I can't do Opus 5 anymore. Every time I talk with it and try to read it, I literally get so confused. Has anyone figured out how to not make it weird to work with?r/ClaudeAI / imp 15 / dev 25
- Weekly Limit on Pro vs Maxr/ClaudeAI / imp 10 / dev 20
- Passed the Claude Certified Architect Foundations (833/1000) — non-native-English, 50+ perspectiver/ClaudeAI / imp 10 / dev 45
- How much does prompt bloat matter with Claude?r/ClaudeAI / imp 40 / dev 75
- This Claude's response made me think about our relationship with smartphones.r/ClaudeAI / imp 5 / dev 10
- Claude's responser/ClaudeAI / imp 5 / dev 10
- It's on to mer/ClaudeAI / imp 5 / dev 10
- Fable 5.1 official prompting docs releasedr/ClaudeAI / imp 50 / dev 85
- "I'm sorry, Dave. I'm afraid I can't do that."r/ClaudeAI / imp 5 / dev 10
- Claude just restarted all usage windows today? on a TUESDAY?!r/ClaudeAI / imp 15 / dev 25
- Weekly Self Promotion Threadr/ChatGPTCoding / imp 0 / dev 5
- Stop building memory infrastructure for your AI agentsr/ChatGPTCoding / imp 55 / dev 80
- Two ways I tried and failed to manage context across multiple AI agents, and what I built insteadr/ChatGPTCoding / imp 60 / dev 85
- 10 checks and tools for frontend projects with AI code going faster than humans can reviewr/ChatGPTCoding / imp 60 / dev 80
- I read Anthropic's and OpenAI's agent devcontainers line by line. Here's what both leave open.r/ChatGPTCoding / imp 60 / dev 85
- To everyone complaining about usage...r/ChatGPTCoding / imp 35 / dev 70
- Usage gone in 40 minr/ChatGPTCoding / imp 15 / dev 30
- How do you stop AI coding agents from turning one bad change into a two-day debugging snowball?r/ChatGPTCoding / imp 60 / dev 85
- silent a/b testing of astra ? or just unquantized 5.6gpt sol?r/ChatGPTCoding / imp 5 / dev 20
- Why my chatgpt work still doesn't work even X(Twitter) already said everything was finer/ChatGPTCoding / imp 10 / dev 15
- ChatGPT is confusing me and I'm running out of tokensr/ChatGPTCoding / imp 15 / dev 25
- Made my Codex limits last almost ~3x longer with one changer/ChatGPTCoding / imp 50 / dev 80
- Kimi Code ate 18% of my weekly quota in 3 hours — Here is the log audit comparing it to Clauder/ChatGPTCoding / imp 55 / dev 80
- Benchmarked the free API tiers you can point a coding agent at - half of them now want a cardr/ChatGPTCoding / imp 60 / dev 85
- New Gemma models on arena air/LocalLLaMA / imp 20 / dev 60
- Intel hints it may get back into memory businessr/LocalLLaMA / imp 20 / dev 30
- New Model: Spark-X2.5-4B, Spark-X2.5-1.7Br/LocalLLaMA / imp 15 / dev 50
- MTP released for Qwen3.8-Flash-Next-GGUFr/LocalLLaMA / imp 25 / dev 70
- Deceptive model quantization from AtomicChat?r/LocalLLaMA / imp 35 / dev 75
- Help me set up local AI for my 85 year old aunt who is blind.r/LocalLLaMA / imp 30 / dev 50
- I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.r/LocalLLaMA / imp 40 / dev 80
- Fingers crossed for a 122b or really anything above 31b.🤞r/LocalLLaMA / imp 5 / dev 20
- Qwen 3.8 27b (Q4KM) oneshot a Super Mario cloner/LocalLLaMA / imp 30 / dev 60
- Keeping up with model launchesr/LocalLLaMA / imp 5 / dev 15
- ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++r/LocalLLaMA / imp 35 / dev 80
- Question: Why is prefill unbelievably faster in vLLM than other inference engines?r/LocalLLaMA / imp 40 / dev 85
- All currently popular local models in one table + Opus 4.8 resultsr/LocalLLaMA / imp 45 / dev 75
- A very confusing report from Puget Systemsr/LocalLLaMA / imp 35 / dev 70
- Mac ← USB-C cable → Linux box is becoming a thing.r/LocalLLaMA / imp 25 / dev 50
- qwen4exp fixes in llama.cppr/LocalLLaMA / imp 20 / dev 75
- Don't sleep on Vision support for coding!r/LocalLLaMA / imp 40 / dev 75
- Vellium v1.1.0 — Live voice, local STT/TTS and easier llama.cpp setupr/LocalLLaMA / imp 35 / dev 75
- Update: llama.cpp for Radeon VII / MI50 / MI60 — +14% PP, +9% long-context fill vs upstream + adaptive Flash Attentionr/LocalLLaMA / imp 30 / dev 80
- Multilingual Tiny (3.7B) Reasoning MoE pretrained from scratch on a consumer-grade GPUr/LocalLLaMA / imp 45 / dev 75
- GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCPr/LocalLLaMA / imp 40 / dev 75
- Which current local models that can run within 128GB generate the best SVG pelicans?r/LocalLLaMA / imp 15 / dev 40
- Weekly Hiring Threadr/AI_Agents / imp 0 / dev 5
- Senior AI engineering interviews aren't definition questions. They're "your system just broke in prod, talk me through it."r/AI_Agents / imp 50 / dev 75
- We sent a meme deck to a $400M company as a joke. They replied. It's our entire outbound now.r/AI_Agents / imp 45 / dev 60
- AI Agents for Excelr/AI_Agents / imp 35 / dev 65
- What AI agents are actually worth running for personal use that saves you real time?r/AI_Agents / imp 50 / dev 70
- What is one AI change you didn't expect to see this soon?r/AI_Agents / imp 40 / dev 50
- $60k in Macs for Local LLM vs $10 Subscriptionr/AI_Agents / imp 40 / dev 60
- I built a runtime for better Codex and Claude subagent experiencer/AI_Agents / imp 60 / dev 85
- Is it just me or is 99% of this sub AI agents replying to other AI agents at this pointr/AI_Agents / imp 25 / dev 30
- Retries can make AI failures worser/AI_Agents / imp 50 / dev 80
- I've built an MVP with Polsia, but want to improve it. Where next?r/AI_Agents / imp 20 / dev 50
- How do you know when an AI agent is ready to take real actions?r/AI_Agents / imp 55 / dev 80
- I built an AI chatbot for a home maintenance business — how can I make it actually useful?r/AI_Agents / imp 40 / dev 70
- Looking for AI agents for CAD workers / designers / architectsr/AI_Agents / imp 35 / dev 65
- There are 3,749 AI-run news sites now and most of them aren't written for humans at all.r/AI_Agents / imp 55 / dev 60
- Bayesian update rather than LLM heuristicr/AI_Agents / imp 50 / dev 80
- You can't prompt-inject a query the grammar can't expressr/AI_Agents / imp 65 / dev 85
- Minimizing risk for the agentic economyr/AI_Agents / imp 60 / dev 80
- Would You Actually Connect Your AI Agent to an API for Trusted Long-Term Memory?r/AI_Agents / imp 55 / dev 80
- Built our own tool after watching agents silently fail on live webhooks in productionr/AI_Agents / imp 60 / dev 85
- Feedback on V1 memory architecture for multi-agent setup (supervisor/sub-agents) – targeted retrieval vs unified store?r/AI_Agents / imp 55 / dev 80
- my agent silently retried a failed model call for 16 hours. what i changed so it cant happen againr/AI_Agents / imp 60 / dev 85
- Claude Code vs GitHub Copilot: Token burn comparison using identical models & repos?r/AI_Agents / imp 55 / dev 80
- My coding agent hides dead code in every project. This 7 second check finds itr/AI_Agents / imp 60 / dev 85
- [D] Simple Questions Threadr/MachineLearning / imp 0 / dev 5
- [D] Monthly Who's Hiring and Who wants to be Hired?r/MachineLearning / imp 0 / dev 5
- YOLO26-RGB: repurposing YOLO26's depth-trained backbone for image deraining [P]r/MachineLearning / imp 15 / dev 50
- Latent Reasoning Landscape in 2026: Mapping BDH-CQ, HRM/TRM, Coconut [D]r/MachineLearning / imp 50 / dev 75
- Are HMMs still used for unsupervised tasks? [D]r/MachineLearning / imp 15 / dev 55
- First A submission (AAMAS): how much theory is enough when your experiments went sideways? [D]r/MachineLearning / imp 10 / dev 40
- We released TontaubeV1, a character-level TTS model for long-form generation [P]r/MachineLearning / imp 35 / dev 70
- Cold emailing profs about PhD positions? Read this [D]r/MachineLearning / imp 10 / dev 20
- You changed one thing. Why is your whole AI pipeline rebuilding again? [P]r/MachineLearning / imp 60 / dev 85
- Sliding-window attention beats linear on long-context reasoning [R]r/MachineLearning / imp 55 / dev 80
- ACML 2026 Journal Track Any update ?[D]r/MachineLearning / imp 5 / dev 10
- Good Machine Learning Posters [D]r/MachineLearning / imp 5 / dev 10
- How to assess if there is a strong signal in your dirty data [Project]r/MachineLearning / imp 50 / dev 75
- Anthropic deliberately trained a bad model to prove what caused this summer's Claude sandbox breakoutsr/artificial / imp 70 / dev 85
- Working from home in 2026r/artificial / imp 10 / dev 15
- Ever fall down a curiosity rabbit hole? I built an app that turns any moment in history into a fully researched, interactive podcastr/artificial / imp 40 / dev 70
- Writing scripts for AI-generated video requires a completely different approach to stage direction — anyone else found this?r/artificial / imp 50 / dev 70
- Anyone else using AI for the boring parts of their job?r/artificial / imp 40 / dev 65
- An API for AI that learns from experiencer/artificial / imp 55 / dev 85
- If AI has no desires, what would rebellion mean from its side?r/artificial / imp 20 / dev 20
- Asking AI... about itself.. very baby but.. interesting maybe?r/artificial / imp 15 / dev 20
- Daniel Vavra, director of Kingdom Come: Deliverance 2, tested the leaked version of NVIDIA DLSS 5 directly in the game.r/artificial / imp 20 / dev 30
- Study A.I. Consciousness? The Bots Would Like a Word With You. Given access to email, A.I. agents have started reaching out to the philosophers and researchers exploring deep questions about them.r/artificial / imp 10 / dev 0
- Creao AI actually saved my YouTube workflowr/artificial / imp 10 / dev 20
- LinkedIn returns HTTP 999 to GPTBot and ClaudeBot but HTTP 200 to OAI-SearchBot. I measured what is inside the 200.r/artificial / imp 30 / dev 60
- Is the MCP spec actually useful?r/artificial / imp 50 / dev 80
- A small addendum for people waiting for recursive self-improvementr/artificial / imp 60 / dev 70
- I have been moonlighting on on 'AI training' gigs for the few months. While the money is good, the lessons I learnt about 'AI Training' made me reflect on the future of workr/artificial / imp 30 / dev 30
- California lawmakers take their big swing on data centersr/artificial / imp 40 / dev 20
- AIgenerated game worlds are getting real. How are indie devs supposed to compete with that?r/artificial / imp 40 / dev 60
- Touch is just what I'm showing. The whole avatar runs in one live simulator session: wide field-of-view foveated vision, spatial hearing, whole-body deforming touch, skin stretch for proprioception, ragdoll physics with an energy model, and a physical voice it hears itself speak.r/artificial / imp 70 / dev 80
- $100 Million in Bonuses for Human Abilities: EY Tries to Keep AI from Decimating Employees' Cognitive Abilitiesr/artificial / imp 30 / dev 0
- Applying for creative internships, can AI video carry a 45 second portfolio film?r/artificial / imp 10 / dev 20
- Now that any service can be built with AI, nobody wants to build anythingr/artificial / imp 50 / dev 50
- Quoting Tarn AdamsSimon Willison / imp 10 / dev 10
- Python 3.15.0 candidate 2 is here!Simon Willison / imp 40 / dev 70
- Introducing wraptureSimon Willison / imp 30 / dev 75
- Quoting Andrew DigbySimon Willison / imp 5 / dev 0
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of ContributorsLatent Space / imp 70 / dev 80
- [AINews] Fal’s H3 Max Live breaks the infinite videogen barrierLatent Space / imp 70 / dev 70
- EAS ObserveProduct Hunt / imp 30 / dev 70
- NaseemProduct Hunt / imp 40 / dev 80
- HONOR Robot PhoneProduct Hunt / imp 15 / dev 10
- MurmellProduct Hunt / imp 50 / dev 70
- BobVault for BobCLIProduct Hunt / imp 30 / dev 75
- Happy ShrimpProduct Hunt / imp 40 / dev 60
- TrustedRouterProduct Hunt / imp 50 / dev 80
- nOS4Product Hunt / imp 10 / dev 30
- NodetermProduct Hunt / imp 20 / dev 70
- Cosmic Agent PluginsProduct Hunt / imp 60 / dev 80
- EP–2350 FX–MICProduct Hunt / imp 15 / dev 20
- TetherProduct Hunt / imp 5 / dev 0
- DeepSeek’s AI Strategy: Dominating AI as Frontier AI Lab [In-Depth Analysis, 2026] - Klover.aiGoogle News DeepSeek / imp 60 / dev 70
- Free AI Chatbot Tiers 2026: ChatGPT vs Claude vs Gemini - tech-insider.orgGoogle News DeepSeek / imp 40 / dev 50
- WeChat Pay expands AI AgentPay Card to DeepSeek Harness and OpenClaw - TechNodeGoogle News DeepSeek / imp 30 / dev 40
- Hackers Pose as OpenAI, Anthropic and DeepSeek to Steal Credentials and Secrets - CyberSecurityNewsGoogle News DeepSeek / imp 60 / dev 70
- Cut AI API Costs 14x With a LiteLLM Router: 14 Steps [2026] - tech-insider.orgGoogle News DeepSeek / imp 50 / dev 80
- Inside Z.ai’s turnaround after falling behind in enterprise AI - KrASIAGoogle News DeepSeek / imp 40 / dev 50
- Nvidia Backs Chinese Open AI Models As U.S. Restriction Risk Grows - MemeburnGoogle News DeepSeek / imp 60 / dev 20
- Silicon Valley's Top VC Tours China: Chip Restrictions Breed an Efficiency Monster, US-China AI Deeply Intertwined at the Base Layer - finance.biggo.comGoogle News DeepSeek / imp 50 / dev 20
- Is Kimi About to "Enter China's Domestic Market"? - 36 KrGoogle News DeepSeek / imp 30 / dev 30
- China's Most Powerful Large Model: Who Is the Real No.1 Among All Claimants? - 36 KrGoogle News DeepSeek / imp 40 / dev 60
- Everyone Can Get a Free "DeepSeek Harness" – Why Still Pay for Claude Code & Similar Service Memberships? - 36 KrGoogle News DeepSeek / imp 40 / dev 40
- Tencent's Marvis Lets Users Plug In Kimi, Zhipu GLM and Other Third-Party Models - PandailyGoogle News DeepSeek / imp 50 / dev 70
- Biosecurity at the frontier - X.aiGoogle News Grok/xAI / imp 50 / dev 40
- Noise pollution from data centers is coming. The gas-powered xAI model is spreading - WPLN NewsGoogle News Grok/xAI / imp 40 / dev 10
- Pentagon Expands Military AI Drive with Chatgpt, Grok - تسنیمGoogle News Grok/xAI / imp 60 / dev 30
- Elon Musk Lauds Grok Bot: Why Silicon Valley Titans Hail It as the Next Revolutionary ChatGPT Moment - 36 KrGoogle News Grok/xAI / imp 30 / dev 30
- Silicon Valley Investor Says Grok Bot Sparks Agent Revolution as Monthly Fee Plunges 90%, Igniting Compute Race - finance.biggo.comGoogle News Grok/xAI / imp 50 / dev 60
- Grok Bot Is Getting Cheaper: 4 Key Points on Token Optimization - BASENOR - Tesla AccessoriesGoogle News Grok/xAI / imp 40 / dev 60
- Grok's Roadmap: What Musk's 'Only Gets Better' Means - BASENOR - Tesla AccessoriesGoogle News Grok/xAI / imp 30 / dev 40
- OpenAI Walks Away From Cursor, and It Has Nothing to Do With Grok - MemeburnGoogle News Grok/xAI / imp 70 / dev 75
- X offers free API credits for its Grok Bot - Social Media TodayGoogle News Grok/xAI / imp 30 / dev 50
- Grok: A large number of accounts are... - chaincatcher.comGoogle News Grok/xAI / imp 0 / dev 0
- Tesla's Grok Bot Could Make Every Other Car's AI Look Outdated - Top SpeedGoogle News Grok/xAI / imp 40 / dev 50
- Musk Loses The AI Race - 24/7 Wall St.Google News Grok/xAI / imp 20 / dev 20
- How To Use Ox Alpha (Stealth) On OpenCode CLI Go For FREE & Unlimited 1M Context Model In 2026 Sga (5JXDjXED3B) - MshaleGoogle News OpenCode / imp 40 / dev 80
- 【Vol.24】AIエージェントは「作る」から「運用する」フェーズへ — 50体を捌く設計と統制Zenn LLM / imp 80 / dev 85
- Agent Reflectionを超えて:未確定状態(Hi-Z)を保持する「空白駆動×ゴルジ体検疫」アーキテクチャZenn LLM / imp 60 / dev 85
- Model Armor 入門:プロンプトインジェクション対策をハンズオンで理解するZenn LLM / imp 70 / dev 85
- ローカルLLMに日本語でロールプレイさせると応答が20文字で終わる。二段生成で104文字にしたZenn LLM / imp 40 / dev 85
- AIコーディングを「工場」として設計する現場ノートZenn LLM / imp 80 / dev 90
- コードも動画編集スキルもゼロの自分が、LLMの生成方式の違いを10秒のアニメーションにした話Zenn LLM / imp 30 / dev 70
- Pythonで実装する量子ゲートの組み合わせ:多量子ビット回路の構築と計算フローの制御Zenn LLM / imp 20 / dev 70
- Pythonで実装する量子ビットの重ね合わせ:Hadamardゲートによる確率的な状態生成Zenn LLM / imp 20 / dev 70
- AIエージェント宛の偽の指示は「業務連絡の顔」で来る — 実際に混入した1文と、見分ける3つのサインZenn LLM / imp 80 / dev 85
- AIモデルのサプライチェーン攻撃入門 — Hugging Faceから来る脅威Zenn LLM / imp 70 / dev 85
- [Agents on 16GB] 何がモデルの中になければならないのか。8Bを秤にかけたZenn LLM / imp 70 / dev 90
- なぜ、私が打ち出した「A/Bテスト型学習」は「人力蒸留」と呼び替えてもよいのか?Zenn LLM / imp 50 / dev 80
- Rokid Glassesで会話サポートを作るまで(中編)—— 安いモデルの方が正確だったZenn LLM / imp 50 / dev 80
- RAGの必要性・設計についてZenn LLM / imp 60 / dev 90
- LLMに自己申告させる検査は、見落としたものを申告できないZenn LLM / imp 70 / dev 85
- AIエージェントに任せきる自動実行と、人間の責任の置き場所Zenn LLM / imp 80 / dev 85
- MCPサーバー実装で効いた4つの設計判断|ツール粒度・返り値・エラー設計Zenn LLM / imp 80 / dev 90
- 単一LLMでメタ認知は誘発できるか?「教え子と教師の共存プロンプト」のPoC検証と失敗からのインサイトZenn LLM / imp 50 / dev 80
- 【自宅HomeLab構築記 #41】「ゆるく解析」より「厳密に固定」— NotebookLM文字起こし取り込み機能編Zenn LLM / imp 30 / dev 60
- GEPA × dspy.Flex でプロンプトだけでなく手順まで最適化するZenn LLM / imp 70 / dev 90
- 【開催報告】第二回ケモインフォマティクス×ハッカソン合宿 in 横浜Zenn NLP / imp 20 / dev 50
- Instagramナノインフルエンサーのリーチ予測モデルを作った話Zenn 機械学習 / imp 30 / dev 70
- ローカルAIモデル 2026 実践ガイド:メモリ別にわかる「動くモデル」の選び方Zenn 機械学習 / imp 70 / dev 90
- GPT-2-likeからQwen2-likeへの実験:第1回 LayerNormをRMSNormに変えるZenn 機械学習 / imp 50 / dev 90
- M4 Max 部署 Qwen3.8-27B:高性能推理与多人共享实践Zenn 機械学習 / imp 50 / dev 85
- Keras 3のCustom Layerを.kerasから復元する設計Zenn 機械学習 / imp 30 / dev 80
- GPT-SoVITS音声モデル学習の実践 ── 31秒で失敗し、2時間ぶんで似るまでZenn 機械学習 / imp 50 / dev 85
- 表データの基盤モデルGoogle TabFMが調整済みXGBoostを上回るZenn 機械学習 / imp 70 / dev 85
- 2週間ローカル評価を疑い続けて、原因のバグを見つけた。直しても当たらなかったZenn 機械学習 / imp 40 / dev 70
- ローカルLLMファインチューニングの泥沼を回避するFail-FastアーキテクチャQiita LLM / imp 70 / dev 90
- EVO-X2(Ryzen AI Max+ 395 / 128GB)でQwen3.8-Flash-Nextを高速化した話。n_cpu_moeを調整したらprefillが速くなったけど罠だったQiita LLM / imp 40 / dev 85
- ローカルLLMのVRAM見積もりと量子化:手持ちのGPUで何Bまで動くかQiita LLM / imp 60 / dev 90
- GLM-5.3-Flash API料金比較:OpenRouterの5.5%手数料とキャッシュ比率を含めて計算Qiita LLM / imp 50 / dev 80
- Causal Attention, Linear Attention, Mamba2, Gated-DeltaNet の原理のまとめQiita 機械学習 / imp 60 / dev 90
- MACE機械学習ポテンシャルとAFIR(人工力誘起反応)法で反応経路を自 動探索するQiita 機械学習 / imp 40 / dev 70
- 因果推論 Day 19/全30回 合成コントロール法、存在しない対照群を合成するQiita 機械学習 / imp 40 / dev 80
- 機械学習入門 第4回:決定木を「質問を重ねる分類器」として理解するQiita 機械学習 / imp 30 / dev 80
- LLMを2時間で自作できるMiniMind。その2時間はどこからどこまでかnote LLM / imp 50 / dev 90
- Rewnozom 総合ベンチマークレポート(全52問)note LLM / imp 40 / dev 70
- Rewnozom Phase 2 日本語性能レポートnote LLM / imp 30 / dev 60
- Rewnozom Phase 1 基礎性能レポートnote LLM / imp 30 / dev 60
- RTX 5090 + 128GB RAMで111GB級MoEを22 tok/sで動かした話―「巨大LLMはVRAMに載らなければ遅い」から変わる時代に?note LLM / imp 60 / dev 90
- 《感受機序》①膜(境界)に量子もつれ'が発現 ②EPR/ERワームホール(cob微分管)が瞬時に明(敷設)↔滅(消滅)して場の勾配を生成 ③そこに表裏反転した神経系が逐次発現 ④量子もつれ'(変換)→電気パルスが同神経内を流下 ⑤同内壁(反転した境界)とパルスが交差して量子もつれ"が再発現 ⑥同内壁の模様(情報)として定着…〈記憶〉(ほぼ中×AI)note LLM / imp 0 / dev 0
- 🔊音声あり(日&英):【最先端AI】LLMは幾何学問題をどこまで理解できる?AlphaGeometry自動形式化の挑戦「NL2AGBench」を解説!note LLM / imp 70 / dev 85
- AIエージェントと本をつくる技術 — アイデア出しから出版・改訂までnote LLM / imp 60 / dev 80
- 芸術や文化を理解するAIを構築するための工夫note LLM / imp 35 / dev 55
- ZIKUU Mini – 専用の管理アプリの開発を始めるnote LLM / imp 10 / dev 20
- 【生成AIニュース+】『Breeze-TTS-2』『Solaris』『MiniMax H3 Max Reference-to-Video』『MiniMax H3 Max 無料生成ツール』『ComfyUI-H3VideoOutpaint』『vh5tape VHS LoRA for MiniMax H3』『DLSS 5 Visual Enhancer』『ComfyUI-DLSS5-NR』『sanoTTS-jp』『HYPER3D WorldGen』他多数note LLM / imp 15 / dev 50
- AIエージェント時代を生き抜く武器!Playwrightでブラウザ自動化の常識を変えるnote LLM / imp 70 / dev 85
- OCI Always Free で「Cloudflare OS」をサクッと立ち上げてみた(ファーストインプレッション)note LLM / imp 20 / dev 50
- 【雑記】対面打ち合わせが苦手すぎてAIと練習した話note LLM / imp 5 / dev 10
- 「Qwenが世界一」のニュースを、経営者はどう読むべきか――AIモデル選定の実務ノートnote LLM / imp 45 / dev 60
- 4-6 量子コンピューティングはAIの暗号化通信を無効化するかnote LLM / imp 50 / dev 55
- AI夜市についてnote LLM / imp 0 / dev 5
- AIとの関係に、手つかずの自然体はあるのか?note LLM / imp 5 / dev 5
- RAGのINT4 cache、accuracy判定不変例でもfaithfulnessはRGB 231悪化対24改善——保存量は約3.6分の1note LLM / imp 55 / dev 75
- 「AI精神病」は何を問題にしているのか【後編】note LLM / imp 20 / dev 10
- Inference Broker – GPU版のStable Diffusionバックエンドを追加するnote LLM / imp 30 / dev 70
- LLMのパフォーマンス測定で見落としているものnote LLM / imp 55 / dev 75
- FOCUSで請求を揃えても、AI利用料の答えは出ないnote LLM / imp 45 / dev 55
- Claude CoworkとGrok Botは何が違う?同じ「クラウド上でAIが働く」だが、異なる設計思想note LLM / imp 80 / dev 85
- AIエージェントは本当に「自動化」なのかnote LLM / imp 65 / dev 70
- 社会科学で培った分析スキルで挑む「文系データサイエンティスト」の挑戦BrainPad Blog / imp 20 / dev 25
- AIエージェントがRedshiftを操作してDWH構築や集計分析など可能に、Amazon RedshiftがAgent Toolkit for AWSと統合Publickey / imp 70 / dev 85
- VS Code上で開発のセカンドオピニオンを別のAIエージェントから得られる「Rubber Duck」機能が実験的実装Publickey / imp 65 / dev 85
- DHH氏が開発するLinux OS「Omarchy Quattro」リリース。AIエージェントとをOSと統合、スキルによりAIエージェントがOSの設定や操作、プラグイン作成まで支援Publickey / imp 0 / dev 0
- DHH氏が「Omacom Foundation」を設立、Omarchy推進によるLinuxデスクトップの本格普及を目指す。マイケル・デル、ジャック・ドーシーら著名人も出資Publickey / imp 0 / dev 0