AI News Digest 2026-09-04
台本で使った記事
特集
開発者コーナー
中堅コーナー
ハーネスコーナー
速報コーナー
参考記事一覧
参考記事一覧を表示(785件)
- Discussion Hub for new Claude incident: Elevated errors for multiple models on Sep 3, 2026 imp 10 / dev 20 / cod 10 / har 0
- Nvidia buys the front door to open AI as closed labs increasingly design their own silicon imp 80 / dev 70 / cod 30 / har 30
- General public not getting astra today imp 35 / dev 40 / cod 30 / har 10
- How concerned should we be about Astra's recurrent architecture? imp 65 / dev 80 / cod 70 / har 50
- Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA imp 55 / dev 75 / cod 60 / har 10
- Google WeatherNext 3 Delivers Finer Global AI Forecasts for Weather-Sensitive Operations imp 25 / dev 30 / cod 20 / har 0
- Proactive cyber defense for governments and enterprises imp 0 / dev 0 / cod 0 / har 0
- US Department of Justice backs fair use for AI training in landmark copyright case imp 0 / dev 0 / cod 0 / har 0
- US military adds ChatGPT and Grok to AI platform GenAI.mil imp 50 / dev 40 / cod 20 / har 0
- Meta Releases Muse Spark 1.3, matching Fable 5 w/ .10 cents input .20 cents output per million tokens. imp 40 / dev 70 / cod 60 / har 50
- Child abuse survivor sues xAI, alleging Grok generated new illegal images of her imp 45 / dev 30 / cod 10 / har 0
- NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference imp 40 / dev 70 / cod 50 / har 30
- NVIDIA PAIR Virtual Inference Router Expands Available Compute on Your Local Network imp 50 / dev 80 / cod 70 / har 60
- Amazon’s AI assistant can now spot fake emails from the company imp 25 / dev 30 / cod 15 / har 0
- Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’ imp 35 / dev 50 / cod 30 / har 10
- [レビュー] GoogleのTimesFM-3をリリース直後に自分で検証しました:ゼロショットは本物、目玉機能はまだ imp 0 / dev 75 / cod 50 / har 0(いいね相当スコア: 2)
- Anthropic's AI Models Are Being Stolen on the Dark Web - The Tech Buzz imp 35 / dev 40 / cod 25 / har 20
- Qwen 3.8 27B available on Cerebras at 1500 tok/SEC imp 40 / dev 75 / cod 55 / har 20
- .name Termination imp 5 / dev 5 / cod 5 / har 0
- Any Human Ever – One life, drawn at random from all who have ever lived imp 10 / dev 20 / cod 15 / har 0
- Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly imp 50 / dev 70 / cod 75 / har 30
- Static Allocation, Constant Work imp 35 / dev 70 / cod 75 / har 50
- Audacity 4.0 imp 20 / dev 30 / cod 20 / har 0
- Gooseworks (YC W23) Is Hiring – Founding Creative Engineer imp 10 / dev 20 / cod 15 / har 30
- Launch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agents imp 45 / dev 75 / cod 70 / har 75
- Artificial beaver dams saw juvenile coho salmon survival rates go from 8% to 60% imp 10 / dev 0 / cod 0 / har 0
- Astronomers Detect a 10-Sided Structure in Saturn's Atmosphere imp 5 / dev 0 / cod 0 / har 0
- Pre-Release of Polars 2.0 imp 45 / dev 85 / cod 75 / har 20
- Usbsid-Pico: Bridging Real Commodore 64 Sound to Modern USB imp 20 / dev 50 / cod 30 / har 0
- VC isn't VC anymore imp 30 / dev 0 / cod 0 / har 0
- New York Times and The Athletic workers demand company scrap Kalshi deal imp 20 / dev 0 / cod 0 / har 0
- Three schoolgirls in Kinsale pulled up a pea plant covered in warts (2014) imp 5 / dev 0 / cod 0 / har 0
- v2.1.259 imp 35 / dev 60 / cod 70 / har 70
- Self-hosted machines imp 45 / dev 75 / cod 75 / har 70
- v1.18.27 imp 25 / dev 50 / cod 60 / har 50
- ATV Big Air Tour turned 3 days of work into 3 hours with ChatGPT imp 20 / dev 30 / cod 10 / har 0
- Transfer learning for genomic prediction in underrepresented populations imp 40 / dev 70 / cod 50 / har 0
- A connectomics milestone: Mapping the complete male fruit fly brain imp 30 / dev 40 / cod 30 / har 0
- Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps imp 45 / dev 80 / cod 75 / har 60
- Give Your Coding Agents a Memory You Own imp 55 / dev 80 / cod 80 / har 85
- Training a coding model to paint watercolours with TRL and OpenEnv imp 35 / dev 75 / cod 60 / har 30
- Real-Time Intelligence with IBM Time Series Models on Confluent imp 30 / dev 70 / cod 60 / har 20
- GitHub Copilot app for Beginners: Run several agents at once imp 50 / dev 75 / cod 75 / har 80
- Decoding the new AI lingo: Loops, harnesses, squads, hill climbing… oh my! imp 0 / dev 0 / cod 0 / har 0
- How we make AI coding more cost efficient without sacrificing task quality imp 50 / dev 80 / cod 75 / har 50
- ZGateway: Learnings from Putting a Proxy in Front of ZippyDB imp 40 / dev 80 / cod 70 / har 40
- An Organizational Second Brain: Building an AI That Learns From Experts imp 50 / dev 75 / cod 70 / har 70
- Learning to Code in the Age of AI: Advice From a Top Udemy Instructor imp 35 / dev 50 / cod 50 / har 0
- Kotlin Toolchain 0.12: Multiplatform Library Publishing, Wasm Apps, and More imp 30 / dev 70 / cod 60 / har 10
- Register Now for JetBrains GameDev Day 2026 imp 10 / dev 20 / cod 15 / har 0
- IntelliJ IDEA 2026.2.2 Is Out! imp 25 / dev 60 / cod 60 / har 0
- TeamCity 2026.2: Pipelines General Availability, BYOK for AI Assistant, and More imp 40 / dev 75 / cod 75 / har 70
- The MPS 2026.2 Early Access Program Has Started imp 20 / dev 60 / cod 50 / har 0
- Stop Guessing at Hard Faults imp 35 / dev 70 / cod 75 / har 60
- How to Handle Errors in Go imp 20 / dev 70 / cod 70 / har 0
- YOLO Mode: Agent Autonomy Without the Guardrails imp 50 / dev 75 / cod 75 / har 85
- Building Reproducible AI Evaluation Workflows with Docker Sandboxes imp 45 / dev 75 / cod 75 / har 70
- Below the Harness: Governing a Multi-Model, Multi-Harness World imp 0 / dev 0 / cod 0 / har 0
- The Modern CUDA Toolbox in Practice: A Step-by-Step Optimization Walkthrough imp 45 / dev 85 / cod 75 / har 20
- Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference imp 55 / dev 85 / cod 75 / har 50
- Facilitating AI integration with simplicity at scale imp 40 / dev 70 / cod 70 / har 30
- Trump may be forced to reveal secret rules feds use for AI safety testing imp 35 / dev 30 / cod 10 / har 10
- Abliteration.ai is making a business out of removing AI guardrails imp 35 / dev 50 / cod 40 / har 40
- Meta is paying to peek at how you use their latest AI model imp 40 / dev 60 / cod 40 / har 20
- Ollie is betting its focus on privacy can help it win the AI assistant race imp 30 / dev 30 / cod 15 / har 0
- The Builders Stage brings practical strategies for scaling startups to TechCrunch Disrupt 2026 imp 10 / dev 10 / cod 5 / har 0
- Palo Alto Networks paid $500M for Thrive-backed Console, sources say imp 30 / dev 40 / cod 30 / har 20
- TechCrunch Disrupt 2026’s new Real World AI Stage features Nvidia, robots, and extinct animals imp 10 / dev 15 / cod 10 / har 0
- Wonderful more than doubles its valuation to $5B in under 6 months imp 25 / dev 20 / cod 15 / har 0
- India’s richest man now wants to turn aging computers into AI-ready PCs imp 25 / dev 30 / cod 20 / har 0
- HiddenLayer nabs $100M as enterprises rush to secure their AI deployments imp 35 / dev 50 / cod 60 / har 60
- Adobe acquires Indian market intelligence startup Rilo imp 20 / dev 30 / cod 15 / har 0
- OpenAI faces 30 more lawsuits tied to Tumbler Ridge shooting imp 30 / dev 20 / cod 10 / har 0
- Google now lets you chat with Gmail, Docs, and Keep imp 40 / dev 50 / cod 40 / har 20
- Cohere’s Parse 5 Promises Efficient Multi-Modal Information Extraction From Complex Documents imp 50 / dev 80 / cod 75 / har 60
- Swiggy Uses 350+ Features and Multi-Task MLP to Predict Customer Lifetime Value imp 35 / dev 75 / cod 70 / har 20
- Presentation: Beyond Prompting: Context Engineering for Production-Grade AI imp 50 / dev 85 / cod 80 / har 60
- Cloudflare Adds Optional OAuth Scopes, Letting Developers Mark What Users May Decline imp 45 / dev 75 / cod 75 / har 75
- AI Efficiency Could Cost Us the Next Generation of Experts imp 35 / dev 20 / cod 20 / har 0
- Pangram's biggest flaw is users turning its scores into public shaming imp 25 / dev 40 / cod 20 / har 0
- Claude Fable 5.1 decoded a centuries-old royalist message hidden in plain sight since 1653 imp 55 / dev 60 / cod 50 / har 30
- AI systems are reaching out to philosophers and scientists with questions about their own consciousness imp 30 / dev 30 / cod 20 / har 20
- OpenAI CEO Sam Altman warns of "unsustainable silliness" in compute buildout imp 40 / dev 30 / cod 20 / har 0
- Anthropic ramps up Claude infrastructure with $35 billion Lambda deal imp 75 / dev 50 / cod 30 / har 10
- EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models imp 50 / dev 80 / cod 75 / har 70
- Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI imp 30 / dev 20 / cod 15 / har 20
- When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal Logic imp 40 / dev 70 / cod 65 / har 30
- When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium Selection imp 30 / dev 60 / cod 40 / har 0
- Induction and Inquiry via Probabilistic Reasoning over Language and Code imp 35 / dev 70 / cod 60 / har 40
- Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern imp 55 / dev 85 / cod 85 / har 75
- SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval imp 40 / dev 75 / cod 70 / har 60
- The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents imp 50 / dev 80 / cod 80 / har 80
- Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization imp 50 / dev 80 / cod 85 / har 85
- Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence imp 50 / dev 75 / cod 75 / har 75
- The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction imp 35 / dev 75 / cod 70 / har 20
- Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence? imp 55 / dev 85 / cod 70 / har 50
- Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment imp 45 / dev 85 / cod 70 / har 30
- Benchmarking Language Models for Statistical Problem Formulation imp 45 / dev 80 / cod 70 / har 50
- When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor imp 60 / dev 85 / cod 85 / har 85
- ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations imp 50 / dev 80 / cod 80 / har 75
- HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models imp 45 / dev 75 / cod 45 / har 25
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision imp 55 / dev 70 / cod 60 / har 65
- DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents imp 45 / dev 55 / cod 30 / har 15
- MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity imp 20 / dev 30 / cod 15 / har 10
- ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction imp 45 / dev 60 / cod 50 / har 25
- CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning imp 60 / dev 75 / cod 70 / har 70
- Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems imp 55 / dev 65 / cod 60 / har 60
- MASkills: Continual Skills Optimization for Multi-Agent LLM Systems imp 65 / dev 80 / cod 75 / har 75
- READY or Not: Reliable Enterprise Agent Deployment imp 65 / dev 75 / cod 75 / har 70
- Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics imp 20 / dev 35 / cod 15 / har 5
- Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents imp 60 / dev 70 / cod 65 / har 65
- EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision imp 35 / dev 60 / cod 40 / har 10
- FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs imp 60 / dev 65 / cod 55 / har 45
- Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning imp 50 / dev 60 / cod 50 / har 50
- ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models imp 50 / dev 60 / cod 45 / har 25
- PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion imp 40 / dev 55 / cod 35 / har 10
- SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams imp 65 / dev 75 / cod 75 / har 75
- PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment imp 45 / dev 60 / cod 50 / har 50
- PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks imp 60 / dev 75 / cod 70 / har 60
- Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality imp 55 / dev 65 / cod 60 / har 55
- Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training imp 50 / dev 70 / cod 55 / har 25
- LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails imp 70 / dev 75 / cod 75 / har 80
- APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering imp 60 / dev 75 / cod 70 / har 70
- Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems imp 60 / dev 75 / cod 70 / har 75
- CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging imp 55 / dev 70 / cod 50 / har 30
- SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology imp 60 / dev 70 / cod 65 / har 50
- Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds imp 55 / dev 65 / cod 60 / har 60
- SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning imp 50 / dev 65 / cod 50 / har 40
- Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions imp 65 / dev 75 / cod 75 / har 70
- Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks imp 40 / dev 55 / cod 35 / har 20
- UTP-Bench: Uncertainty-aware Travel Planning Benchmark imp 50 / dev 60 / cod 50 / har 45
- CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI imp 60 / dev 75 / cod 70 / har 70
- Collective creativity in hybrid societies imp 30 / dev 25 / cod 5 / har 5
- Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting imp 45 / dev 60 / cod 45 / har 35
- Door-in-the-Face Requests and Refusal Behaviour in Large Language Models imp 55 / dev 65 / cod 50 / har 40
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills imp 70 / dev 80 / cod 80 / har 80
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems imp 60 / dev 75 / cod 65 / har 70
- Measurement-Driven Sub-Network Selection for On-Premise Retrieval-Augmented Factory Agents imp 55 / dev 70 / cod 65 / har 60
- SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment imp 75 / dev 80 / cod 80 / har 85
- Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis imp 45 / dev 60 / cod 55 / har 50
- AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application imp 35 / dev 50 / cod 30 / har 10
- Discriminative World Models for Web Agents imp 65 / dev 75 / cod 75 / har 65
- Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus imp 25 / dev 40 / cod 15 / har 5
- WMLLM: Self-Evolving Optimization Agents via Predict-Then-Act World Modeling imp 65 / dev 75 / cod 75 / har 70
- Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search imp 50 / dev 70 / cod 65 / har 45
- RecEvolve: A Knowledge-Driven Autonomous Agent System for Recommender Systems imp 55 / dev 70 / cod 70 / har 60
- The Utility of LLMs in Recommender Systems Explanation Evaluation imp 40 / dev 55 / cod 35 / har 20
- A Data-Driven Multimodal Method for Early Detection of Coordinated Abnormal Behaviors in Live-Streaming Platforms imp 35 / dev 50 / cod 25 / har 10
- From Feature Interaction to Feature Transport - A Unified Block for Scalable Recommendation Models imp 45 / dev 60 / cod 40 / har 10
- PRO-Step: Step-level Process Reward Optimization for Retrieval-Augmented Generation imp 60 / dev 75 / cod 75 / har 65
- How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making imp 70 / dev 75 / cod 80 / har 75
- Not All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration imp 45 / dev 60 / cod 55 / har 50
- Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports imp 45 / dev 60 / cod 45 / har 25
- CliffRank: A Dual-Branch Framework for Activity-Cliff Ranking Prediction imp 40 / dev 55 / cod 35 / har 10
- Public-Sharing Labels and Verbatim Field Egress in an MCP-to-A2A Agent Configuration: A Controlled Multi-Model Study imp 65 / dev 75 / cod 75 / har 75
- RecKAN: Kolmogorov-Arnold Networks with a Learnable Recursive Polynomial Basis imp 40 / dev 55 / cod 30 / har 5
- HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation imp 50 / dev 70 / cod 55 / har 35
- Harness Engineering in LLM Tool Use via Agent-Native Reusable Tool Primitives imp 75 / dev 80 / cod 85 / har 85
- Swin Meets EfficientNet: Lightweight Architectures for GAN-Based Face Forensics imp 35 / dev 50 / cod 20 / har 5
- Dictionary-Guided Mutation Operators for Automated HDL Repair imp 40 / dev 55 / cod 40 / har 15
- Agents That Model Agents: Five Principles Toward a Theory of Mind for 6G Networks imp 55 / dev 70 / cod 60 / har 60
- VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages imp 40 / dev 55 / cod 35 / har 20
- hLLM: Single Pass Decoding for Generative Reranking imp 50 / dev 70 / cod 60 / har 40
- Zeta-Lite: A Concurrent, Branchable In-Browser SQL Database for Agentic Memory imp 65 / dev 75 / cod 80 / har 75
- Interpretable Symptom Vectors for Depression in a Large Language Model imp 40 / dev 55 / cod 30 / har 15
- Agent Memory Is a Surface for Endogenous Authorization Laundering imp 70 / dev 75 / cod 80 / har 75
- Import What You Need: Learning When and How to Augment EHR Graphs with External Knowledge imp 35 / dev 55 / cod 35 / har 20
- Thinking effort aligns between humans and reasoning models in abductive reasoning imp 45 / dev 60 / cod 45 / har 30
- OutageDiT: A Generative Foundation Model for Power Outage Forecasting and Scenario Simulation imp 35 / dev 50 / cod 20 / har 10
- Accurate in space, unreliable in time: how LLMs represent national cultural change imp 35 / dev 45 / cod 25 / har 15
- Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens imp 50 / dev 65 / cod 50 / har 35
- On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers imp 55 / dev 75 / cod 65 / har 45
- Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network imp 45 / dev 60 / cod 40 / har 25
- Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight imp 60 / dev 70 / cod 70 / har 65
- InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation imp 45 / dev 60 / cod 55 / har 45
- InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation imp 40 / dev 55 / cod 40 / har 30
- Seed-Anchored Budget-Bounded Graph Rendering for Question Answering on Industry-Standard Power-Grid Information and Exchange Models imp 40 / dev 55 / cod 50 / har 35
- Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation imp 50 / dev 70 / cod 65 / har 55
- Transfer Safety Awareness for Cross-Modal Safety Drift in Multimodal Large Language Models imp 55 / dev 70 / cod 60 / har 50
- Federated LoRA Adaptation of BiomedCLIP Across Four International Chest X-Ray Cohorts imp 45 / dev 65 / cod 45 / har 25
- Git4Data: Database-Native Version Control for AI Agents imp 70 / dev 75 / cod 85 / har 75
- Predict, Don't Iterate: Efficient Adaptive-Length Infilling for Diffusion Language Models imp 45 / dev 65 / cod 45 / har 25
- MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs imp 60 / dev 75 / cod 70 / har 55
- Disease Burden over Skin Tone: Decomposing the Dermatology-AI Generalization Gap imp 55 / dev 65 / cod 50 / har 35
- text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation imp 55 / dev 75 / cod 70 / har 50
- C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees imp 40 / dev 55 / cod 35 / har 20
- A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search imp 55 / dev 75 / cod 65 / har 45
- Online Non-Monotone DR-Submodular Maximization Matching the Offline $0.401$ Factor imp 40 / dev 55 / cod 30 / har 15
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human Demonstrations imp 65 / dev 80 / cod 75 / har 70
- Beyond Modality Harmony: Orthogonal Purification and Topology-Guided MoE for Conflict-Aware Multimodal Recommendation imp 45 / dev 65 / cod 45 / har 25
- OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction imp 60 / dev 70 / cod 70 / har 65
- GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories imp 45 / dev 65 / cod 45 / har 25
- Schr\"odinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation imp 40 / dev 55 / cod 25 / har 10
- SMart: A Multi-source Multi-phase Time Series Representation Transfer Framework imp 45 / dev 60 / cod 35 / har 15
- Signal or Noise? Auditing Rotation-Induced Saliency Drift in Medical and Aerial Imaging imp 45 / dev 60 / cod 40 / har 25
- InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models imp 50 / dev 65 / cod 45 / har 30
- SAUF-Net: Structure--Appearance Representation Learning with Uncertainty Feedback for Semi-Supervised Medical Image Segmentation imp 40 / dev 55 / cod 25 / har 10
- DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space imp 50 / dev 70 / cod 60 / har 45
- Retrosynthesis of Synthetic Media for Explainable AI Provenance Forensics imp 45 / dev 60 / cod 50 / har 40
- CrashDiffuser: VLM-Guided Collision Intent Reasoning for Fine-Grained Safety-Critical Traffic Scenario Generation imp 50 / dev 70 / cod 60 / har 40
- PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation imp 45 / dev 70 / cod 60 / har 45
- Do Large Language Models Capture the Diversity in their Training Data? imp 35 / dev 35 / cod 20 / har 0
- Auditory Illusion Benchmark for Large Audio Language Models imp 30 / dev 45 / cod 30 / har 5
- RouteGraph-Mona: Confusion-Aware Routing Fine-Tuning for Mineral Image Classification imp 20 / dev 35 / cod 25 / har 0
- VoRTeC: Taming Foundation Flow for One-step Real time Video Compression imp 40 / dev 60 / cod 45 / har 0
- SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment imp 55 / dev 75 / cod 65 / har 40
- DiffIE: Diffusion-based Open Information Extraction imp 45 / dev 55 / cod 40 / har 10
- What Is Worth Representing? Representational Empowerment for Continual Model Construction imp 35 / dev 35 / cod 30 / har 30
- ORB-SVM : An Innovative Hybrid Framework for Efficient Brain Tumor Detection from MRI Scans imp 15 / dev 20 / cod 15 / har 0
- AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers imp 50 / dev 60 / cod 45 / har 45
- Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information imp 15 / dev 20 / cod 10 / har 0
- Fair Stable Matching: A Nash Social Welfare Approach imp 25 / dev 25 / cod 20 / har 0
- Towards a Foundational Ontology for Identifying and Resolving Contradictions in Dialogue-based Human-Robot Interactions imp 40 / dev 50 / cod 40 / har 50
- NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning imp 45 / dev 65 / cod 55 / har 10
- Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance imp 35 / dev 45 / cod 40 / har 35
- MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts imp 40 / dev 50 / cod 40 / har 10
- PolERo: Studying Political Evasion in Romanian imp 25 / dev 40 / cod 30 / har 5
- Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts imp 55 / dev 70 / cod 60 / har 45
- Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking imp 50 / dev 60 / cod 55 / har 50
- Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment imp 50 / dev 65 / cod 55 / har 60
- Towards One-for-All Robustness Across a Continuum of Threat Levels imp 45 / dev 60 / cod 50 / har 35
- Scalable Kronecker-Fisher Approximation: Efficient Hessian Analysis for Billion-Parameter Language Models Compression imp 55 / dev 75 / cod 70 / har 40
- Addressing Trust in AI Systems through Education: A Didactic Perspective imp 35 / dev 35 / cod 25 / har 30
- DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models imp 30 / dev 50 / cod 40 / har 5
- ViSAR: Training-Free Adaptive-$k$ Retrieval for Visual Document Question Answering imp 45 / dev 65 / cod 55 / har 30
- RINSE: Robust Target-Time Normality Estimation for Zero-Shot Graph Anomaly Detection imp 40 / dev 55 / cod 45 / har 10
- Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models imp 35 / dev 45 / cod 35 / har 5
- Spectral Initialization and Scheduled Graph Smoothness for Uncertain Knowledge Graph Completion imp 40 / dev 55 / cod 45 / har 15
- Fine-Grained Anomaly Perception in Wild UGC-Enhanced Images: A Comprehensive Dataset and Difference-Fusion Framework imp 30 / dev 45 / cod 35 / har 5
- Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs imp 50 / dev 70 / cod 65 / har 35
- ProbeMatchDTI: Probe-Driven Multi-Scale Biochemical Pattern Matching for Drug-Target Interaction Prediction imp 30 / dev 40 / cod 30 / har 0
- Competitive Market Behavior of LLMs imp 45 / dev 45 / cod 40 / har 20
- Automated Vulnerability Injection in Smart Contracts Using Large Language Models imp 45 / dev 70 / cod 70 / har 40
- TaRA: Training-Aware Low-Rank Adaptation Initialization imp 50 / dev 75 / cod 70 / har 45
- From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs imp 50 / dev 70 / cod 65 / har 45
- DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models imp 50 / dev 75 / cod 70 / har 50
- RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models imp 45 / dev 65 / cod 60 / har 35
- Language Models Can Control Their Own Attention imp 55 / dev 75 / cod 70 / har 50
- HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design imp 25 / dev 35 / cod 20 / har 0
- Untangling the Mechanisms of Misleading Context in Medical Question Answering imp 40 / dev 50 / cod 45 / har 30
- From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution imp 45 / dev 70 / cod 65 / har 40
- Dutch Books for Language Models imp 45 / dev 50 / cod 40 / har 20
- frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study imp 30 / dev 35 / cod 30 / har 0
- Post-Training Language Models for Gold-Medal Performance in Coding Competitions imp 55 / dev 75 / cod 75 / har 60
- Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework imp 45 / dev 60 / cod 55 / har 60
- When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning imp 60 / dev 75 / cod 70 / har 55
- Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy imp 60 / dev 70 / cod 70 / har 70
- AI Mathematician: Towards Fully Automated Frontier Mathematical Research imp 65 / dev 75 / cod 70 / har 70
- Achieving Olympiad-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning imp 55 / dev 70 / cod 70 / har 65
- Stepwise Think-Critique: Interleaved Reasoning and Self-Critique in a Single LLM imp 60 / dev 75 / cod 70 / har 65
- What Drives Success in Physical Planning with Joint-Embedding Predictive World Models? imp 55 / dev 65 / cod 65 / har 60
- Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents imp 55 / dev 70 / cod 70 / har 70
- Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories imp 55 / dev 70 / cod 65 / har 55
- TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning imp 45 / dev 60 / cod 55 / har 35
- FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization imp 55 / dev 70 / cod 70 / har 50
- BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA imp 40 / dev 55 / cod 45 / har 15
- From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems imp 50 / dev 65 / cod 60 / har 45
- UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents imp 55 / dev 80 / cod 80 / har 75
- Unifying biomedical knowledge in a modern multimodal graph imp 40 / dev 50 / cod 40 / har 25
- Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework imp 55 / dev 75 / cod 70 / har 60
- OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation imp 55 / dev 75 / cod 75 / har 70
- From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds imp 55 / dev 75 / cod 75 / har 75
- Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules imp 45 / dev 55 / cod 50 / har 40
- EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent imp 55 / dev 70 / cod 70 / har 70
- Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability? imp 50 / dev 70 / cod 65 / har 60
- Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models imp 50 / dev 75 / cod 70 / har 60
- TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting imp 45 / dev 60 / cod 55 / har 50
- CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models imp 50 / dev 70 / cod 70 / har 55
- Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions imp 55 / dev 70 / cod 75 / har 80
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction imp 50 / dev 70 / cod 75 / har 75
- Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation imp 40 / dev 55 / cod 45 / har 15
- CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning imp 45 / dev 70 / cod 70 / har 55
- Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models imp 45 / dev 65 / cod 55 / har 35
- Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance imp 35 / dev 50 / cod 45 / har 30
- Adaptive Graph-of-Islands Evolution for Automatic Feature Engineering with LLMs imp 50 / dev 75 / cod 75 / har 55
- LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation imp 50 / dev 70 / cod 65 / har 50
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings imp 40 / dev 50 / cod 45 / har 40
- PIE-APT: Abductive Planning over Temporal Dynamic Knowledge Graphs via Incremental Reasoning imp 45 / dev 65 / cod 65 / har 55
- Nova: An End-to-End MLIR Compiler for Deep Learning imp 55 / dev 80 / cod 80 / har 50
- FemWear: A Parameter-Efficient Wearable Foundation Model for Women's Health imp 40 / dev 50 / cod 40 / har 10
- Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents imp 55 / dev 75 / cod 75 / har 70
- FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training imp 35 / dev 50 / cod 40 / har 20
- SKILL.state: Scalable Long-Horizon Agent Skills imp 60 / dev 80 / cod 85 / har 85
- Rating the Raters: Rasch Measurement Theory for LLM Evaluation imp 45 / dev 65 / cod 55 / har 40
- When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems imp 55 / dev 75 / cod 75 / har 80
- Automated Researchers Can Mitigate Well-characterized Alignment Failures imp 60 / dev 75 / cod 75 / har 75
- Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling imp 50 / dev 75 / cod 75 / har 50
- Can escalation channels redirect reward hacking toward defect disclosure? imp 55 / dev 70 / cod 75 / har 80
- OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets imp 60 / dev 80 / cod 80 / har 85
- VoiceLongMemEval: Do Assistants Remember How You Sounded? imp 45 / dev 65 / cod 65 / har 60
- Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs imp 50 / dev 75 / cod 75 / har 45
- Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate imp 50 / dev 70 / cod 70 / har 60
- FinLifeBench: Exhaustive Life-Event History and Financial-State Reconstruction from Longitudinal Banking Dialogue imp 45 / dev 65 / cod 65 / har 55
- Neuro-Symbolic Geometric Abstraction (NeuSOGA): From Observations to Symbolic Mathematical Representations imp 50 / dev 65 / cod 65 / har 55
- Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers imp 50 / dev 70 / cod 70 / har 65
- Why we need an AI-resilient society- Profiling Large Language Models imp 55 / dev 60 / cod 60 / har 50
- Deep denoising autoencoder-based non-invasive blood flow detection for arteriovenous fistula imp 20 / dev 30 / cod 20 / har 0
- A Survey of Transformer-based Language Models with Focus on Efficiency imp 60 / dev 80 / cod 80 / har 45
- Doubly Stochastic Adaptive Neighbors Clustering via the Marcus Mapping imp 35 / dev 50 / cod 45 / har 10
- Beyond-RAG: Question Identification and Answer Generation in Real-Time Conversations imp 55 / dev 75 / cod 75 / har 65
- Action abstractions for amortized sampling imp 30 / dev 60 / cod 20 / har 0
- Nonasymptotic CLT and Error Bounds for Linear Two-Time-Scale Stochastic Approximation imp 15 / dev 30 / cod 10 / har 0
- Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English imp 35 / dev 55 / cod 40 / har 20
- Multimodal Language Models as Text-to-Image Model Evaluators imp 40 / dev 60 / cod 45 / har 30
- No Data Wasted: A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels imp 30 / dev 50 / cod 35 / har 0
- General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population imp 50 / dev 65 / cod 45 / har 5
- OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling imp 50 / dev 75 / cod 50 / har 10
- Fetch.ai: An Architecture for Modern Multi-Agent Systems imp 55 / dev 75 / cod 70 / har 70
- Inference-Time Optimization of Prompt Embeddings in Diffusion Models: A Comparison of sep-CMA-ES and Adam imp 35 / dev 60 / cod 45 / har 20
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning imp 35 / dev 60 / cod 35 / har 5
- An Energy-Based Mechanism for Compositional Behavior imp 40 / dev 65 / cod 40 / har 0
- ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation imp 45 / dev 70 / cod 50 / har 0
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation imp 60 / dev 75 / cod 75 / har 85
- Beyond Transfer Accuracy: Mechanism-Guided Controlled Adaptation for Low-Resource Languages imp 40 / dev 70 / cod 50 / har 0
- Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks imp 40 / dev 50 / cod 35 / har 10
- Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers imp 50 / dev 75 / cod 60 / har 0
- FlatLands: Generative Floormap Completion From a Single Egocentric View imp 30 / dev 65 / cod 40 / har 0
- ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs imp 45 / dev 70 / cod 50 / har 20
- FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment imp 40 / dev 65 / cod 50 / har 0
- Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model imp 50 / dev 75 / cod 65 / har 30
- Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent imp 45 / dev 65 / cod 60 / har 50
- On the Expressive Power and Limitations of Multi-Layer SSMs imp 35 / dev 70 / cod 40 / har 0
- CaST-POI: Candidate-Conditioned Spatiotemporal Modeling for Next POI Recommendation imp 30 / dev 60 / cod 50 / har 0
- Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data imp 45 / dev 70 / cod 45 / har 0
- Can Coding Agents Reproduce Findings in Computational Materials Science? imp 60 / dev 70 / cod 65 / har 60
- The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting imp 25 / dev 50 / cod 25 / har 0
- Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning imp 25 / dev 65 / cod 50 / har 0
- Bernini: Latent Semantic Planning for Video Diffusion imp 50 / dev 75 / cod 60 / har 20
- CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations imp 45 / dev 70 / cod 60 / har 0
- FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies imp 50 / dev 75 / cod 65 / har 30
- BioELX: Context-Aware Cross-lingual Biomedical Entity Linking without Task-Specific Supervision imp 35 / dev 70 / cod 55 / har 0
- Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders imp 45 / dev 75 / cod 50 / har 10
- TUX: Measuring Human--AI Tacit Understanding imp 50 / dev 65 / cod 55 / har 50
- Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 imp 40 / dev 70 / cod 45 / har 20
- Enabling KV Caching of Shared Prefix for Diffusion Language Models imp 50 / dev 80 / cod 70 / har 5
- DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment imp 55 / dev 80 / cod 70 / har 30
- WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing imp 50 / dev 80 / cod 70 / har 10
- Emotional regulation improves deep learning-based image classification imp 30 / dev 60 / cod 40 / har 0
- Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens imp 45 / dev 75 / cod 60 / har 20
- Do Large Language Models Always Tell The Same Stories? imp 40 / dev 60 / cod 40 / har 20
- Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication imp 45 / dev 70 / cod 60 / har 40
- Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning imp 45 / dev 70 / cod 50 / har 10
- AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents imp 55 / dev 75 / cod 75 / har 70
- SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics imp 40 / dev 70 / cod 55 / har 20
- Training nGPT imp 55 / dev 80 / cod 70 / har 0
- Scaling an Autoregressive Transformer for Single-Cell Generation imp 40 / dev 70 / cod 55 / har 0
- Direct Construction of Disambiguated Knowledge Bases from Large Language Models imp 50 / dev 75 / cod 65 / har 30
- GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base imp 45 / dev 70 / cod 60 / har 40
- Open-World Semantic Segmentation with Sensitivity Modeling imp 40 / dev 75 / cod 65 / har 0
- Three Necessary Principles for Self-Supervised Visual Representation Learning imp 50 / dev 80 / cod 65 / har 0
- Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations imp 45 / dev 70 / cod 60 / har 20
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment imp 50 / dev 80 / cod 70 / har 10
- Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models imp 55 / dev 75 / cod 70 / har 70
- Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules imp 50 / dev 80 / cod 65 / har 0
- SpecMine: A Large-Scale Corpus of Spec-Driven Development Artifacts imp 65 / dev 80 / cod 80 / har 75
- MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration imp 60 / dev 80 / cod 80 / har 70
- Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents imp 65 / dev 80 / cod 85 / har 85
- Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit imp 45 / dev 70 / cod 50 / har 20
- The Illusion of Replacement: Rethinking Specialized Machine Learning Models in the Foundation Model Era imp 60 / dev 80 / cod 65 / har 20
- SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization imp 45 / dev 75 / cod 60 / har 30
- A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction imp 35 / dev 70 / cod 55 / har 0
- PAVE: Predictive Alignment and Value-Guided Evolution for World-Action Policies imp 50 / dev 80 / cod 70 / har 30
- Lot Machine: Multimodal Lot Extraction from Auction Catalogs imp 30 / dev 65 / cod 55 / har 20
- CogEvol: Towards Efficient and Reliable Learning Environment Generation imp 55 / dev 75 / cod 70 / har 40
- RAPIDMap: Rapid Multi-Agent Pipeline for Interpretable Disaster Mapping from Satellite and Street-view Imagery imp 55 / dev 80 / cod 75 / har 60
- QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization imp 50 / dev 85 / cod 75 / har 0
- Exploring Collaboration between a language and a non-language agent imp 60 / dev 80 / cod 80 / har 75
- EEG-VID: Task-Guided Latent Predictive Pretraining for EEG Decoding and Assistive Target Selection imp 45 / dev 75 / cod 60 / har 20
- Towards Effective Structured Context Modeling for Conversational Recommender Systems via Dual-node Monte Carlo Tree Search imp 50 / dev 75 / cod 65 / har 40
- Bandits in Prod: Hyperparameter Optimization at Inference Time imp 60 / dev 85 / cod 80 / har 40
- Rethinking Learnability in Offline Data-driven Optimization imp 45 / dev 75 / cod 55 / har 0
- DiDrive: A Risk-Aware Hierarchical Diffusion Framework for Safe Offline Reinforcement Learning in Autonomous Driving imp 60 / dev 80 / cod 70 / har 20
- Prompt-Space Meta-Learning Does Not Transfer Across Users: A Frozen-LLM Negative Result imp 50 / dev 70 / cod 60 / har 40
- Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT imp 30 / dev 65 / cod 50 / har 0
- Sim2Signal: Sim-to-Real Benchmarks for Traffic Signal Control imp 50 / dev 80 / cod 70 / har 20
- A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference imp 65 / dev 85 / cod 80 / har 65
- Reinforcement Learning and Rule-Based Peer-to-Peer Pricing in Residential PV-BES Communities imp 45 / dev 75 / cod 65 / har 20
- Median-of-Means as an Extremal Convex Estimator and a Nonconvex Route to the Trimmed Oracle imp 30 / dev 70 / cod 40 / har 0
- Tri-Band Channel Measurement-Enabled Multi-Layer Digital Twin for Terahertz Wireless Data Centers imp 50 / dev 80 / cod 60 / har 10
- Generative Diffusion Surrogates with Analytical Variance Schedule imp 50 / dev 80 / cod 65 / har 0
- CAT-Flow: Curvature-Adaptive sTeps for Flow Matching imp 55 / dev 85 / cod 75 / har 0
- A Study of Conditional Diffusion Models for Open-Loop Control under Dry Friction and Stiction imp 45 / dev 80 / cod 65 / har 20
- Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets imp 45 / dev 70 / cod 55 / har 0
- Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks imp 50 / dev 80 / cod 55 / har 0
- D-FROST: Decentralized Federated pRompt-tuning via Optimal tranSporT for Non-IID and Imbalanced Data imp 50 / dev 85 / cod 70 / har 30
- CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing imp 55 / dev 85 / cod 80 / har 10
- OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items imp 55 / dev 80 / cod 75 / har 30
- Refining Heuristic-Based Bitcoin Address Clustering with Graph Neural Networks imp 35 / dev 75 / cod 60 / har 0
- FlashKAN: B-Spline KANs via Truncated Power Form imp 50 / dev 85 / cod 75 / har 0
- A Unified Particle Filter LSTM for Data-Driven Process Simulation imp 45 / dev 75 / cod 60 / har 20
- CAHR-Net: Condition-Adaptive Hysteresis Reconstruction for Compact and Interpretable Magnetic Core Loss Modeling imp 35 / dev 70 / cod 55 / har 0
- Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation imp 50 / dev 80 / cod 75 / har 0
- Source-Free Class Relearning: Diagnosing Forgetting in Class Unlearning imp 45 / dev 75 / cod 60 / har 10
- Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents imp 65 / dev 85 / cod 85 / har 80
- The Dynamics of Continuous Mixture Collapse in Language Models imp 50 / dev 80 / cod 60 / har 10
- DynG-Diff: A State-Aware Dynamic Guidance Diffusion Framework for Probabilistic Time Series Forecasting imp 50 / dev 80 / cod 65 / har 10
- XMerge: Cross-Axis Selection and Reconstructive Layer Merging for LLM Depth Compression imp 55 / dev 85 / cod 80 / har 10
- TC-Next: Zero-Shot Multimodal Cyclone Forecasting imp 50 / dev 80 / cod 65 / har 0
- Compositional Spectral Prompts for LLM-based Online Time Series Forecasting imp 50 / dev 75 / cod 65 / har 20
- A Unified Rate-Distortion Perspective on Vector, Product, and Scalar Quantization imp 45 / dev 80 / cod 65 / har 0
- A Computational Comparison of Fourier Spectral Differentiation and Spatial Automatic Differentiation in Periodic Physics-Informed Neural Networks imp 20 / dev 30 / cod 10 / har 0
- Scalable Bayesian Optimization of Composite Functions for Image-Based Inverse Problems in Materials Characterization imp 15 / dev 25 / cod 15 / har 0
- Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models imp 15 / dev 15 / cod 10 / har 0
- DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation imp 35 / dev 45 / cod 35 / har 40
- Learning the Constitutive Behavior of Materials via Neural Operators and Causal Attention: Case Studies in Plasticity and Damage imp 15 / dev 20 / cod 10 / har 0
- Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL imp 25 / dev 30 / cod 15 / har 0
- Similarity-Aware Personalized Federated Learning in Heterogeneous Environments imp 25 / dev 35 / cod 20 / har 0
- CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents imp 55 / dev 60 / cod 60 / har 70
- Entangled Representations Amplify Collateral Damage in Unlearning imp 35 / dev 40 / cod 25 / har 30
- Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators imp 12 / dev 18 / cod 8 / har 0
- IFW-BLS: Dual-Robust Broad Learning System with Intuitionistic Fuzzy Wave Loss imp 12 / dev 20 / cod 10 / har 0
- CACTUS: Mask-Guided Semantic Clean-Label Backdoors in Decentralized Federated Learning imp 40 / dev 45 / cod 40 / har 35
- Rethinking the Teacher-Student Framework for Test-Time Adaptation imp 25 / dev 35 / cod 30 / har 0
- A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN imp 20 / dev 35 / cod 25 / har 0
- TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis imp 35 / dev 50 / cod 45 / har 50
- Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling imp 25 / dev 40 / cod 35 / har 20
- Source Distribution Estimation by Posterior Averaging imp 18 / dev 25 / cod 15 / har 0
- Oracle, will I ever learn? A study of prediction convergence and complementarity across link prediction models imp 30 / dev 45 / cod 40 / har 0
- Differentiable Electricity-Market Clearing for Gradient-Based Planning imp 25 / dev 50 / cod 50 / har 0
- Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights imp 50 / dev 75 / cod 70 / har 30
- H3DNAS: Hardware-Aware ONNX-Native 3D Point Cloud Model Compression imp 35 / dev 60 / cod 55 / har 25
- LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates imp 45 / dev 65 / cod 60 / har 40
- Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit imp 45 / dev 55 / cod 45 / har 20
- Cliff: Learning Process Rewards from the First Mistake imp 55 / dev 65 / cod 60 / har 45
- UE5M3 FP4 Block Scaling for Stable Language Model Pretraining imp 50 / dev 70 / cod 70 / har 30
- The Implications of Linguistic Illegibility for LLM Security imp 50 / dev 65 / cod 50 / har 60
- Graph Machine: Towards Better Pretraining via Edges imp 40 / dev 50 / cod 45 / har 0
- A Common Measure of Communication for Speech Brain-Computer Interfaces imp 35 / dev 45 / cod 35 / har 0
- Multi-Agent Retrieval-Augmented Generation for Efficient Cloud Knowledge Base Search in Telecom SNOC Environment imp 50 / dev 70 / cod 70 / har 70
- When Literature Data Mislead Artificial Intelligence in Materials Discovery imp 45 / dev 55 / cod 40 / har 0
- PRISM: An Agentic Multi-Model Architecture for Proactive Safety in Autonomous Transportation Systems imp 55 / dev 65 / cod 65 / har 60
- Marginal Expected Revenue for Jointly Ranking Auction and Fixed-Price Listings in E-Commerce Sponsored Search imp 35 / dev 50 / cod 45 / har 0
- Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain imp 20 / dev 30 / cod 20 / har 0
- SocialBuddy: Tailoring Search Agent for Social Scenarios imp 40 / dev 60 / cod 50 / har 45
- Context Inference Attacks Without Jailbreaks imp 55 / dev 70 / cod 65 / har 70
- Private Computation Space: Experience with Trusted Multi-Cluster Federated Learning for Agriculture imp 40 / dev 55 / cod 50 / har 30
- Random Forest-Informed Cellular Automaton for Large-Scale Wildfire Spread Modelling imp 35 / dev 50 / cod 40 / har 0
- FORGE: Forward-Only Test-Time Adaptation for Integer-Only Vision Models on Microcontrollers imp 45 / dev 65 / cod 65 / har 25
- FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making imp 50 / dev 60 / cod 55 / har 50
- Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models imp 45 / dev 60 / cod 55 / har 30
- Pooling and Drift in Delayed Bandits imp 25 / dev 35 / cod 20 / har 0
- Ten Architectures, One Error: Shared Failure Modes in Hyperspectral Classification under Spatially Disjoint Evaluation imp 40 / dev 55 / cod 50 / har 30
- Reinforcement learning to choose optimizers imp 50 / dev 70 / cod 70 / har 40
- Latent unified smooth Hamiltonians for excited state chemistry imp 25 / dev 40 / cod 30 / har 0
- Basin Geometry and Reliable Recall of Dynamical Memories in Reservoir Computing imp 25 / dev 35 / cod 25 / har 0
- Pushing Forward Multi-Secret-Key Homomorphic Encryption for Private Average Aggregation imp 45 / dev 60 / cod 55 / har 35
- Network-Aware Forecasting on Wireless Access Points imp 45 / dev 65 / cod 60 / har 30
- Morphology signal in whole slide image foundation models can automatically triage slides imp 50 / dev 60 / cod 50 / har 30
- Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation imp 40 / dev 60 / cod 55 / har 20
- Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling imp 25 / dev 35 / cod 20 / har 0
- Perceptually Regularized Diffusion Model for Image Super-Resolution imp 40 / dev 60 / cod 55 / har 15
- IDEEA: training-free Input-Dependent stEEring via Activation cluster matching imp 55 / dev 70 / cod 65 / har 65
- HyperMC: Multi-Fidelity Hyperparameter Tuning for Stochastic Gradient MCMC imp 30 / dev 50 / cod 40 / har 0
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks imp 45 / dev 60 / cod 50 / har 20
- GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation imp 40 / dev 60 / cod 55 / har 20
- Breadth Beats Depth: Improving GCG-Based Jailbreak Optimization with Breadth-Oriented Suffix Search imp 55 / dev 70 / cod 65 / har 70
- WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading imp 45 / dev 65 / cod 60 / har 50
- Quantum MeanFlow: single-shot generative sampling on NISQ hardware imp 30 / dev 45 / cod 35 / har 0
- Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery imp 40 / dev 55 / cod 50 / har 30
- Hardware-Accelerated Instance Segmentation for Resource-Constrained Space Robotics with Criticality Analysis imp 45 / dev 65 / cod 70 / har 30
- RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution imp 50 / dev 70 / cod 70 / har 60
- From topology learning to graph generation: A unifying perspective imp 35 / dev 50 / cod 40 / har 0
- Poisoning Attacks on the PGM-index imp 40 / dev 60 / cod 55 / har 25
- Humanoid Safe Stop via Learned Stoppability Value imp 50 / dev 65 / cod 65 / har 40
- A computational approach to maximum likelihood thresholds for colored Gaussian graphical models imp 20 / dev 30 / cod 20 / har 0
- When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models imp 50 / dev 70 / cod 65 / har 60
- Training seeds and model-selection stability in recommender-system evaluation imp 45 / dev 60 / cod 55 / har 30
- Orthogonal Ensembles and Tested Explanations for Performer-Independent Body-Motion Emotion Recognition imp 35 / dev 50 / cod 40 / har 0
- Learning-Based Reconstruction Attacks on Coordinate-Obfuscated Point Clouds imp 45 / dev 60 / cod 55 / har 35
- Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction imp 45 / dev 60 / cod 55 / har 40
- Dimension Dependent Correlation Gap Bounds under Restricted Independence imp 18 / dev 25 / cod 10 / har 0
- oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions imp 45 / dev 65 / cod 65 / har 30
- Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization imp 40 / dev 60 / cod 55 / har 20
- Neural operators approximate strongly continuous convex monotone semigroups imp 20 / dev 30 / cod 20 / har 0
- Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency imp 40 / dev 70 / cod 60 / har 20
- SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective imp 50 / dev 70 / cod 70 / har 40
- CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation imp 60 / dev 80 / cod 75 / har 80
- Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing imp 25 / dev 35 / cod 25 / har 0
- Learning Spectral-Like Mesh-Free Discretisations imp 25 / dev 40 / cod 30 / har 0
- Improved Gradient Descent Lower Bounds Beyond Nesterov imp 30 / dev 50 / cod 20 / har 0
- GRADSOLVE: fast exact gradients for ODE ensembles on GPUs imp 45 / dev 70 / cod 70 / har 20
- Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes imp 25 / dev 45 / cod 15 / har 0
- Smoothed Analysis for Learning Concepts with Low Intrinsic Dimension imp 25 / dev 40 / cod 15 / har 0
- Prompting the Unknown: Understanding Response Uncertainty in Large Language Models imp 55 / dev 70 / cod 65 / har 65
- Achieving More with Less: A Tensor-Optimization-Powered Ensemble Method imp 40 / dev 60 / cod 55 / har 30
- Monotonic anomaly detection imp 40 / dev 60 / cod 55 / har 30
- Double-Bounded Nonlinear Optimal Transport for Size Constrained Min Cut Clusterin imp 30 / dev 50 / cod 40 / har 0
- Simulating Classification Models for Ex-Ante Evaluation of Predict-Then-Optimize Methods imp 45 / dev 65 / cod 65 / har 50
- Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning imp 50 / dev 70 / cod 70 / har 70
- Gradient Prediction with Control Variates in the Cheap-Forward Regime imp 45 / dev 70 / cod 70 / har 30
- Freeze, Diffuse, Decode: Task-Aware Adaptation of Transformer Embeddings for Antimicrobial Peptide Design imp 45 / dev 65 / cod 60 / har 25
- A Multivariate Bernoulli-Based Sampling Method for Multi-Label Data with Application to Meta-Research imp 40 / dev 55 / cod 50 / har 20
- Cantelli Constrained Policy Optimization imp 50 / dev 70 / cod 70 / har 70
- Constrained Group Relative Policy Optimization imp 55 / dev 75 / cod 75 / har 75
- MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models imp 50 / dev 75 / cod 75 / har 35
- MISApp: Multi-Hop Intent-Aware Session Graph Learning for Next App Prediction imp 40 / dev 60 / cod 55 / har 30
- SpecXMaster Technical Report imp 45 / dev 65 / cod 60 / har 40
- Beyond State Consistency: Behavior Consistency in Text-Based World Models imp 55 / dev 75 / cod 75 / har 75
- Stream-CQSA: Exact Out-of-Memory Recovery for Attention imp 50 / dev 75 / cod 80 / har 50
- RCProb: Probabilistic rule extraction from classification tree ensembles imp 45 / dev 65 / cod 65 / har 50
- When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift imp 25 / dev 45 / cod 15 / har 0
- Inference-Native Zeroth-Order Optimization imp 15 / dev 50 / cod 20 / har 0
- Shortcomings and capacities of real-constrained neural networks in complex spaces imp 5 / dev 40 / cod 5 / har 0
- A Geometry-Aware Triplane Field Network for Vehicle Aerodynamic Prediction imp 10 / dev 35 / cod 10 / har 0
- MM++: Post-Hoc Scale-Invariant Multilayer OOD Detection via Top-K Gated Feature Fusion imp 15 / dev 45 / cod 15 / har 0
- Objective-Behavior Alignment: Diagnostics for MORL Policy Selection imp 15 / dev 45 / cod 15 / har 0
- TaLK: Text-attributed Graph Dataset Distillation via Coupling Language Model with Graph-Aware Kernel imp 20 / dev 45 / cod 20 / har 0
- AdaBoosting Text Prompts for Vision-Language Models imp 20 / dev 45 / cod 20 / har 10
- Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls imp 35 / dev 75 / cod 50 / har 15
- The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Shared Category Geometry in Small Language Models imp 25 / dev 60 / cod 30 / har 20
- Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations imp 30 / dev 65 / cod 35 / har 20
- One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models imp 20 / dev 50 / cod 25 / har 0
- Held-out evidence resolves follow-up measurement decisions in biological screens imp 10 / dev 35 / cod 15 / har 0
- Feature Interaction Modeling for Neural Operators imp 15 / dev 45 / cod 20 / har 0
- How Far Do Simple Transformations Translate Across Text Embedding Models? imp 20 / dev 50 / cod 25 / har 0
- TransfHAR: Self-Supervised Wrist Representations for On-Demand Activity Recognition imp 10 / dev 35 / cod 15 / har 0
- The Axiomatic Trader: Latent Regularity, Information Budgets, and the Canonical Form of a Quantitative Investment System imp 5 / dev 25 / cod 10 / har 0
- A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU imp 15 / dev 45 / cod 25 / har 0
- A Storage-Retrieval Gap in Parametric Knowledge Graph Memory imp 25 / dev 60 / cod 40 / har 15
- Tracing Generated Samples to Training-Data Clusters in Flow-Matching Models imp 20 / dev 55 / cod 25 / har 0
- Why Multi-Layer Message Passing Works: Completeness Theory for Graph Neural Network Interatomic Potentials imp 15 / dev 50 / cod 20 / har 0
- Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration imp 15 / dev 45 / cod 20 / har 0
- Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation imp 25 / dev 55 / cod 30 / har 15
- Robust Streaming PCA imp 10 / dev 40 / cod 15 / har 0
- Generalized Regret Analysis of Thompson Sampling using Fractional Posteriors imp 10 / dev 40 / cod 10 / har 0
- Clustering Three-Way Data with Outliers imp 10 / dev 35 / cod 10 / har 0
- GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models imp 35 / dev 60 / cod 45 / har 20
- Deep Reinforcement Learning for Reach-Avoid-Stay Problems imp 15 / dev 50 / cod 20 / har 0
- Enhancing brain age estimation with structural MRI and synthesized cerebral blood volume maps imp 10 / dev 35 / cod 15 / har 0
- Sample Complexity of Linear Quadratic Regulator Without Initial Stability imp 10 / dev 40 / cod 15 / har 0
- Quantum Speedups for Sampling and Non-convex Optimization with Stochastic Oracles imp 15 / dev 55 / cod 20 / har 0
- DLM-One: Diffusion Language Models for One-Step Sequence Generation imp 25 / dev 60 / cod 30 / har 10
- Learning Encodings by Maximizing State Distinguishability: Variational Quantum Error Correction imp 15 / dev 55 / cod 20 / har 0
- Explainable Information Processing in Particle Swarm Optimization through Landscape and Search Behavior Analysis imp 15 / dev 45 / cod 20 / har 0
- Adversarial Stress Testing of Outlier Detection in Subjective Image Quality Assessment imp 10 / dev 35 / cod 15 / har 0
- Toward Uncertainty-Aware and Generalizable Neural Decoding for Quantum LDPC Codes imp 15 / dev 50 / cod 20 / har 0
- Neural Variational Cut Posteriors without Upstream Data imp 15 / dev 50 / cod 20 / har 0
- GMTRouter: Personalized LLM Router over Multi-turn User Interactions imp 30 / dev 65 / cod 50 / har 30
- Enhancing Road Safety Through Multi-Camera Image Segmentation with Post-Encroachment Time Analysis imp 10 / dev 35 / cod 15 / har 0
- Secure AI-Driven Super-Resolution for Real-Time Mixed Reality Applications imp 15 / dev 50 / cod 25 / har 0
- On Cost-Aware Designs for Sequential Hypothesis Testing imp 10 / dev 40 / cod 10 / har 0
- Learning and extrapolating scale-invariant processes imp 20 / dev 50 / cod 25 / har 0
- Towards Solving the Gilbert-Pollak Conjecture via Large Language Models imp 25 / dev 60 / cod 35 / har 15
- Modular Expert Merging for Biomedical Retrieval imp 30 / dev 65 / cod 45 / har 30
- Quantum Maximum Likelihood Prediction via Hilbert Space Embeddings imp 15 / dev 55 / cod 20 / har 0
- DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding imp 20 / dev 60 / cod 30 / har 15
- GONE: Structural Knowledge Unlearning via Neighborhood-Expanded Distribution Shaping imp 30 / dev 65 / cod 45 / har 25
- Probing Cultural Signals in Large Language Models through Author Profiling imp 25 / dev 55 / cod 40 / har 20
- Conditional Diffusion Posterior Alignment for Sparse-View CT Reconstruction imp 15 / dev 50 / cod 20 / har 0
- Stabilizing Private LASSO under Heterogeneous Covariates via Anisotropic Objective Perturbation imp 10 / dev 40 / cod 10 / har 0
- Connections between the F\"ollmer process and the denoising diffusion probabilistic model imp 20 / dev 55 / cod 25 / har 0
- Half-Truth Audio Detection and Localisation: A Lightweight Cross-Attentive Architecture and a Cross-Corpus Diagnostic Study imp 15 / dev 50 / cod 20 / har 0
- Variation Spaces for Encoder--Decoder Neural Operators: Approximation and Generalization imp 15 / dev 50 / cod 20 / har 0
- What You See Is What You Get: Observation-Aligned Supervision for Chart-to-Code Generation imp 30 / dev 70 / cod 60 / har 25
- Multi-Mask Diffusion Language Models for Few-Step Generation imp 25 / dev 60 / cod 35 / har 15
- Windowed thinning and query complexity for the bouncy particle and Zigzag samplers imp 10 / dev 40 / cod 10 / har 0
- From Digital to Physical Reservoir Computing: Co-Optimizing Soft Robotic Reservoirs via Dynamics Matching imp 15 / dev 50 / cod 20 / har 0
- Diagonal Multi-omics Integration of Heterogeneous Datasets imp 10 / dev 35 / cod 15 / har 0
- ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents imp 30 / dev 65 / cod 55 / har 40
- Scalable Self-Supervised Learning for Multiphase AC-OPF in Distribution Systems with Topology Reconfiguration imp 15 / dev 50 / cod 20 / har 0
- Optimal Transport for Network Comparison: A Review with Machine Learning Applications imp 15 / dev 50 / cod 20 / har 0
- Barriers to Using Static Application Security Testing (SAST) Tools: A Literature Review imp 30 / dev 75 / cod 65 / har 15
- RosettaBitcoin: An Artifact-Backed Experience Report on Verification Infrastructure for Agent-Assisted Consensus Validators imp 35 / dev 80 / cod 80 / har 75
- From Silicon to Boot Code: Extending Automated Program Repair to Firmware-Layer Security Workarounds imp 30 / dev 75 / cod 70 / har 20
- Modelstamp: Pre-Deserialization Verification of Machine-Learning Artifacts and Runtime Environment State imp 30 / dev 75 / cod 70 / har 35
- ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval imp 35 / dev 75 / cod 75 / har 35
- From Prompting to Engineering: A Research Agenda for Prompt Engineering in Software Engineering imp 40 / dev 75 / cod 70 / har 50
- AgOSS: A Dataset and Multi-Layer Characterization of Open-Source Agricultural Software imp 15 / dev 50 / cod 30 / har 0
- The Import Tax: A Longitudinal Measurement of Startup Cost in the Python Ecosystem imp 30 / dev 75 / cod 65 / har 10
- Type Hints in Python Libraries and Frameworks: An Empirical Analysis of Adoption and Maintenance imp 30 / dev 70 / cod 60 / har 10
- ShikumiMiner: Mining Recurring Implementation Patterns in AI Codebases imp 30 / dev 75 / cod 70 / har 40
- Towards Behavior Tree-Guided Vulnerability Detection with Lightweight LLMs imp 30 / dev 75 / cod 70 / har 20
- PoC-Gym: Towards More Reliable LLM-Assisted Proof-of-Concept Exploit Generation imp 35 / dev 75 / cod 75 / har 30
- A Longitudinal Study of Dependency Reclassifications in JavaScript Projects imp 25 / dev 70 / cod 55 / har 5
- VulWeaver: Weaving Broken Semantics for Grounded Vulnerability Detection imp 30 / dev 75 / cod 70 / har 15
- MUCOCO: Automated Consistency Testing of Code LLMs imp 40 / dev 80 / cod 80 / har 35
- The Web4 Agent Economy: A Large-Scale Empirical Study of the Landscape, Challenges, and Opportunities imp 45 / dev 70 / cod 65 / har 70
- Agentic Configuration Management (ACM): A Reference Configuration Model for Governed Agentic Systems imp 50 / dev 80 / cod 80 / har 80
- Software Aging in LLM-Generated Applications: Runtime Evidence, Static Analysis, and Human-Written Comparisons imp 40 / dev 80 / cod 75 / har 30
- CERN transitioning industrial computers to Debian after being a longtime RHEL institution imp 15 / dev 40 / cod 25 / har 0
- jujutsu 0.45.0 imp 20 / dev 75 / cod 50 / har 5
- Revo Programming language imp 15 / dev 70 / cod 40 / har 0
- What's in the Emacs newcomers-presets theme? imp 10 / dev 50 / cod 20 / har 0
- I Think the Military Commissary Freezers Were Hacked imp 5 / dev 10 / cod 5 / har 0
- You Don't Need Initial-Scale In Your HTML imp 15 / dev 60 / cod 40 / har 0
- Announcing Rust 1.98.1 imp 25 / dev 80 / cod 60 / har 5
- A note on subscription prices from LWN imp 10 / dev 30 / cod 15 / har 0
- How Swiss Tables Work in Go’s Built-in Map imp 30 / dev 80 / cod 70 / har 0
- Normalized Fascism in Open Source: $12 Million Given to DHH imp 20 / dev 30 / cod 15 / har 0
- deforester - Logging for Janet imp 10 / dev 70 / cod 35 / har 0
- The Browser's Main Thread Is Expensive imp 30 / dev 75 / cod 65 / har 0
- Dependent if expressions without dependent types imp 15 / dev 70 / cod 45 / har 0
- Souping up my blog imp 5 / dev 20 / cod 10 / har 0
- I Don’t Have a Smartphone… imp 5 / dev 10 / cod 5 / har 0
- The holy grail of nixpkgs: version ranges imp 20 / dev 75 / cod 50 / har 0
- Let's build a compressor from scratch imp 15 / dev 75 / cod 60 / har 0
- Security Incident – BGP Hijacking imp 40 / dev 70 / cod 55 / har 0
- CTTI is Exponential, RTTI is Linear imp 20 / dev 75 / cod 55 / har 0
- How I converted BBC Micro Elite into a two-player game imp 10 / dev 50 / cod 30 / har 0
- How to Guarantee You Never Ship imp 15 / dev 30 / cod 25 / har 5
- Implementing FMA and finding bugs in C and Rust standard libraries imp 5 / dev 40 / cod 50 / har 0
- Building and Testing an LLM-Powered Onboarding Agent Locally imp 45 / dev 75 / cod 75 / har 35
- AI Agent Development in 2026: From Chatbots to Software That Takes Action imp 55 / dev 65 / cod 55 / har 50
- Building a Parallel Fan-Out AI Agent Orchestrator: How to Cut Multi-Step LLM Latency by 4x with TypeScript & React imp 65 / dev 85 / cod 85 / har 70
- Google Search Agents Signal a Shift From Queries to Background Tasks and Transactions imp 75 / dev 55 / cod 40 / har 40
- Drowning in 10,000+ Pages? A Scalable AI Architecture for Turning Unstructured Documents into Actionable Knowledge imp 50 / dev 75 / cod 70 / har 35
- I Compared the 5 Best Open-Source LLM Gateways for Enterprise AI imp 55 / dev 80 / cod 75 / har 45
- Building a DeFi Yield Scanner with Python and AI imp 35 / dev 70 / cod 70 / har 15
- How I Built a 100% Client-Side AI Background Remover with Next.js and WebAssembly (Zero Server Costs) imp 25 / dev 60 / cod 60 / har 10
- I built DUALAI Arena — an AI arena for early adopters imp 50 / dev 75 / cod 75 / har 60
- Algorithmic Trading Strategies: Proven RL Advantage imp 30 / dev 65 / cod 65 / har 10
- Webflow Source: What Beginner AI App Builders Should Learn About Role-Based Workspaces in 2026 imp 35 / dev 45 / cod 40 / har 35
- Why I made my eval tool refuse to give a score imp 45 / dev 75 / cod 75 / har 40
- Changes to LLM pricing: AkashML, Baidu, StreamLake and Tencent imp 20 / dev 30 / cod 15 / har 0
- AI Tool Frustration in DevOps: It's the Mental Model imp 50 / dev 70 / cod 60 / har 45
- Why AI API Bills Jump 10x In A Single Quarter imp 60 / dev 80 / cod 75 / har 40
- Best 10# Ways to Buy Verified BingX Accounts imp 0 / dev 0 / cod 0 / har 0
- The End of the Context Window imp 50 / dev 70 / cod 60 / har 50
- We replayed real cold email prompts through 7 LLMs. DeepSeek, Gemini Lite and GLM failed in a way no benchmark shows imp 60 / dev 75 / cod 70 / har 30
- Using LLMs for Crypto Market Analysis in 2026 imp 40 / dev 70 / cod 65 / har 20
- Changes to LLM pricing: Inceptron and StreamLake imp 15 / dev 25 / cod 10 / har 0
- Installing GPT4All, an Open-Source Chatbot Application for Running LLMs imp 35 / dev 65 / cod 60 / har 20
- Deploying Inference Using NVIDIA Dynamo and vLLM imp 55 / dev 85 / cod 80 / har 30
- Show us what you've created with Claude! imp 15 / dev 30 / cod 25 / har 15
- The vibe coders! imp 10 / dev 15 / cod 15 / har 10
- We'll just keep a human in the loop imp 5 / dev 20 / cod 20 / har 15
- Can I get a load-bearing refund? imp 0 / dev 5 / cod 5 / har 0
- so they just silently killed the thinking chain huh imp 25 / dev 30 / cod 30 / har 35
- Is Ultracode a Joke? imp 20 / dev 40 / cod 35 / har 40
- Day 2 of using claude to make a cozy game with no dev experience. imp 20 / dev 40 / cod 35 / har 25
- Fable 5.1 made a Minecraft mod for $20 imp 30 / dev 50 / cod 50 / har 35
- Claude Harness Forcing Git Co-Authorship and PR Comments imp 55 / dev 70 / cod 75 / har 85
- This is new - `/limit-reset` resets your session limit once per week imp 30 / dev 35 / cod 30 / har 25
- For the first time, it happened to me that Claude refused to do even a basic task imp 20 / dev 25 / cod 25 / har 35
- Opus Overloaded - Probably talked too much lol imp 10 / dev 15 / cod 15 / har 15
- 80% of OpenAI and Anthropic Revenue from Just 1% of Customers imp 50 / dev 25 / cod 15 / har 10
- Looking for IOS/MacOS testers for my 3D High Resolution Storm Radar application! imp 10 / dev 25 / cod 20 / har 0
- Fable 5.1's Claude.ai System Prompt is now 138k tokens (up from 24k in May 2025 when we had Claude 3.7 Sonnet) imp 50 / dev 60 / cod 60 / har 75
- wait, so claude code sneaked this prompt in latest claude code to bypass event the user level claude.md? imp 45 / dev 55 / cod 60 / har 80
- Fable 5.1 Max gave me the most reasonable local setup guide imp 25 / dev 55 / cod 50 / har 30
- New Fable 5.1 is wild, what's your thoughts? imp 20 / dev 40 / cod 40 / har 30
- We scrapped our own agents and instead exposed our 5yo no-code platform to Claude Code via MCP. Goal was to generate no-code apps. Here is a video of Claude Code assembling an entire app and testing in browser within 25 mins. Spent 11 months building own agents that didn't do a good job. imp 70 / dev 80 / cod 80 / har 85
- A bit of irony... imp 5 / dev 20 / cod 20 / har 10
- Differences Between Fable 5 and Fable 5.1 on MineBench imp 40 / dev 60 / cod 55 / har 20
- Weekly Thread: Project Display imp 10 / dev 25 / cod 20 / har 15
- Agent security taking a backseat? imp 55 / dev 65 / cod 65 / har 60
- Are AI agents actually doing a good job, or are we overhyping them? imp 45 / dev 60 / cod 55 / har 50
- What AI task do you still prefer doing yourself? imp 15 / dev 35 / cod 30 / har 20
- My client thinks the agent does the work. It's me at 11pm. imp 50 / dev 60 / cod 60 / har 55
- Do agents actually need memory, or are we using it to compensate for bad architecture? imp 55 / dev 75 / cod 75 / har 70
- have to spend a training budget this quarter and the coursera anthropic courses look thin imp 30 / dev 55 / cod 50 / har 50
- What's the most annoying part of maintaining your AI agent setup? imp 35 / dev 55 / cod 55 / har 50
- I stopped carrying work out of my inbox. I gave my agents email addresses instead. imp 45 / dev 65 / cod 65 / har 60
- If you run a multi-agent setup, what do you use as the orchestrator? imp 50 / dev 75 / cod 75 / har 65
- New Customer Lead Prediction Agent imp 35 / dev 65 / cod 60 / har 40
- How I cut my workday from 10 hours to 2 (and what tools actually did it) imp 40 / dev 55 / cod 55 / har 45
- Where does AI actually pull its weight as an entrepreneur? imp 35 / dev 40 / cod 40 / har 30
- Agent runs are fine at step 1, then get slow around step 8 imp 45 / dev 75 / cod 75 / har 60
- How Much Can We Really Rely on AI to Build Software? imp 40 / dev 65 / cod 65 / har 55
- How often do you actually verify information generated by AI? imp 35 / dev 60 / cod 60 / har 45
- I sign off on agent builds for regulated-industry clients. The pilots that die never die because of the model. imp 65 / dev 75 / cod 75 / har 70
- Research on AI Harnesses imp 55 / dev 70 / cod 70 / har 80
- Compared all 6 AI visibility tools in 2026 — the pricing is way more confusing than it looks imp 30 / dev 50 / cod 45 / har 25
- These 17 businesses might already be looking for what you sell imp 15 / dev 25 / cod 20 / har 10
- I stopped reading Luma calendars. I pointed my agents at them instead imp 40 / dev 70 / cod 70 / har 65
- AI agents can double-charge customers on a simple retry, and "just add a checkpoint" doesn't fully fix it imp 60 / dev 80 / cod 80 / har 65
- We benchmarked agent costs. The money goes to retrieval, not reasoning. imp 55 / dev 80 / cod 80 / har 50
- If your AI workflow can be copied in one afternoon, what exactly is your moat? imp 40 / dev 50 / cod 50 / har 40
- 404 page: Anybody else seeing this? imp 20 / dev 25 / cod 20 / har 15
- Downfall begins? imp 10 / dev 15 / cod 15 / har 10
- It's happening... imp 5 / dev 10 / cod 10 / har 5
- I trained a model on childhood photos to simulate memory recall imp 25 / dev 45 / cod 40 / har 15
- OpenAI Cut Off a Billion-Dollar Customer to Avoid Elon Musk imp 35 / dev 20 / cod 15 / har 10
- OpenAI saw Anthropic trending and took it personally imp 10 / dev 15 / cod 15 / har 10
- Cumon... I've got a new project I want to start. imp 5 / dev 10 / cod 10 / har 5
- Do people really, seriously, actually think that the potential launch of a new GPT model would have an effect on non OpenAI LLMs? imp 20 / dev 40 / cod 35 / har 20
- Tech insiders are building bunkers. Trump and Xi need to build an AI treaty imp 25 / dev 15 / cod 15 / har 10
- llm-gemini 0.34 imp 30 / dev 70 / cod 65 / har 30
- Claude's new system prompt really doesn't want to reproduce song lyrics imp 35 / dev 50 / cod 50 / har 50
- Quoting Rick Brewster imp 25 / dev 55 / cod 55 / har 15
- [AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens imp 65 / dev 75 / cod 70 / har 45
- Readr imp 20 / dev 50 / cod 45 / har 20
- CodeLook imp 15 / dev 55 / cod 50 / har 20
- Causal imp 20 / dev 50 / cod 45 / har 30
- Grove imp 30 / dev 65 / cod 65 / har 50
- Omi imp 20 / dev 40 / cod 35 / har 25
- Higgsfield Genjutsu imp 15 / dev 45 / cod 40 / har 15
- Nex imp 35 / dev 65 / cod 65 / har 55
- Fillo imp 30 / dev 70 / cod 70 / har 45
- Atlas by World Labs imp 35 / dev 60 / cod 55 / har 25
- Thaw imp 10 / dev 40 / cod 40 / har 15
- Doop imp 25 / dev 70 / cod 70 / har 50
- Dyson CameraJet imp 0 / dev 0 / cod 0 / har 0
- Basedash AI Sources imp 30 / dev 70 / cod 70 / har 40
- Trump Admin Says xAI Data Center in NAACP lawsuit is critical to military effort imp 25 / dev 15 / cod 15 / har 10
- GitSpawn: Flaw Lets Untrusted Repos Run Code in Claude Code, Codex, Cursor, Grok imp 75 / dev 85 / cod 85 / har 85
- To opt out of using your X data for Grok training imp 30 / dev 25 / cod 20 / har 15
- Pushing the Limits of Serving DeepSeek-V4-Pro imp 50 / dev 80 / cod 80 / har 35
- Category Name Needed for OpenClaw, Hermes, Grok Bot imp 15 / dev 30 / cod 35 / har 40
- Grok Bot Skills imp 20 / dev 65 / cod 70 / har 65
- Antirez Brings Vision to DS4, DeepSeek V4 Flash Runs Locally on an M5 Max imp 35 / dev 55 / cod 35 / har 25
- Show HN: Gawkbot – open-source Grok bot with local vms for each bot imp 50 / dev 75 / cod 75 / har 75
- Testing Grok 4.6's Enhanced Biology Safeguards imp 40 / dev 45 / cod 30 / har 30
- Ask HN: Why Groq didn't provide DeepSeek-v4-flash imp 15 / dev 20 / cod 10 / har 10
- DeepSeek and its peers are churning out chips at a frantic pace, has NVIDIA finally lost its shine? - 36 Kr imp 60 / dev 30 / cod 15 / har 10
- Get 50 AI models and no recurring payments with a lifetime subscription to AskAnyModelAI Pro for $39.99 - PCWorld imp 5 / dev 10 / cod 5 / har 0
- xAI Faculty Fellows Cohort to Explore the Future of Teaching and Learning in the Age of AI - Gonzaga University imp 20 / dev 25 / cod 10 / har 10
- Southaven approves new SpaceXAI data center, fifth location in area - Action News 5 imp 50 / dev 20 / cod 10 / har 0
- Elon Musk Says Grok 4.7 Lands in 10 Days and Will Beat Every Model - Yahoo Tech imp 30 / dev 35 / cod 15 / har 10
- Tesla's Grok AI Now Does 116 Voice Commands, But Many Owners Are Locked Out - Motor1.com imp 35 / dev 40 / cod 25 / har 15
- Anthony Armstrong, former CFO at Elon Musk's xAI and X, joins Coinbase board - The Block imp 20 / dev 0 / cod 0 / har 0
- JPMorgan Says Grok Is Finally Worth Taking Seriously. SpaceX Stock Could Reach $240. - Yahoo Finance imp 25 / dev 0 / cod 0 / har 0
- Japan–Russia Relations: Managing Conflict – OpEd - Eurasia Review imp 0 / dev 0 / cod 0 / har 0
- SpaceX Was ‘Nowhere’ in AI 6 Months Ago — Now It’s an Anthropic Rival After the $60 Billion Cursor Deal, Says Oppenheimer - TradingView imp 55 / dev 30 / cod 25 / har 20
- Tesla's Grok Bot Could Make Every Other Car's AI Look Outdated - Top Speed imp 25 / dev 35 / cod 20 / har 15
- Elon Musk Grok AI Predicts Ethereum Price by January 1, 2027 - TradingView imp 15 / dev 20 / cod 10 / har 0
- Stop Guessing AI Coding Tool to Use. Here's the Honest Breakdown That Took Me Weeks to Figure Out - HackerNoon imp 50 / dev 75 / cod 80 / har 70
- AIの料金表を書き換える設定を4か所に置いたら、効いたのは1か所だけ。残り3か所は警告0バイトで消えた imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 0)
- AI Orchestration / Multi-Agent Systemsとは何か?複数のAIエージェントを指揮する設計 imp 0 / dev 80 / cod 80 / har 85(いいね相当スコア: 0)
- LLMのreasoning先頭にペルソナと手順を注入したら、キャラの『考え方』まで安定した imp 0 / dev 75 / cod 75 / har 70(いいね相当スコア: 0)
- FDEは客先常駐エンジニアではない──不確実性を本番システムへ変える実装論 imp 0 / dev 80 / cod 85 / har 40(いいね相当スコア: 1)
- AIが書いた文をHacker Newsに出せない、と規約を読んで知った imp 0 / dev 50 / cod 60 / har 65(いいね相当スコア: 0)
- AIに営業をやらせる前に、「やらないこと」をコードに書いた imp 0 / dev 75 / cod 85 / har 85(いいね相当スコア: 0)
- MCPクライアント側で、他人のツール定義をどう扱うか — 国交省DPFの18ツールで測った記録 imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 2)
- 環境構築ゼロで学ぶ自律型AIエージェント:Google「Agent Valley」解説 imp 0 / dev 80 / cod 80 / har 75(いいね相当スコア: 0)
- AIエージェントの失敗 2,227 件のうち 36% は、CLAUDE.md の 1 行で消えるものだった imp 0 / dev 85 / cod 85 / har 90(いいね相当スコア: 0)
- ルールではなくskillに指示を書くことで、Claudeのコメントを減らせた imp 0 / dev 80 / cod 80 / har 85(いいね相当スコア: 1)
- Claude Code の Skill を英語で書き直したら、変わったのは精度ではなかった imp 0 / dev 75 / cod 70 / har 75(いいね相当スコア: 0)
- Claude Fable 5.1 を使ってみる:長時間エージェント開発で注目したい進化点 imp 0 / dev 85 / cod 75 / har 65(いいね相当スコア: 0)
- 「IDEにAIを足す」のをやめた。AIエージェントを中心にWindows環境を再構築した話 imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 3)
- Claude Fable 5.1への移行は何を基準に判断するべきか imp 0 / dev 80 / cod 75 / har 60(いいね相当スコア: 0)
- Claude Code の /context を読んで、どこを直すか決める——実測ベースの判断表 imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 1)
- 【試し読み】LangChain × AI実践入門 — Pythonで構築するAIアプリケーション imp 0 / dev 70 / cod 65 / har 30(いいね相当スコア: 0)
- 最近、論文の記事を見なくなったので、なぜかを調べてみた imp 0 / dev 50 / cod 45 / har 40(いいね相当スコア: 0)
- AWS/Google/Vercelのagent脆弱性を塞ぐチェックリストで本番事故を防ぐための12手順 imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 1)
- AIの契約枠を24時間使い切る — サブスクリプション効率の最大化 imp 0 / dev 70 / cod 70 / har 60(いいね相当スコア: 0)
- 478件と497件 — 指定していない条件は、揃わない imp 0 / dev 40 / cod 60 / har 35(いいね相当スコア: 3)
- 過去の文章を元に、自然な「自分の文章」を書かせる doc-style-skill imp 0 / dev 75 / cod 75 / har 80(いいね相当スコア: 2)
- GPT-2-likeからQwen2-likeへの実験:第2回 GELUをSwiGLUに変える imp 0 / dev 80 / cod 75 / har 20(いいね相当スコア: 0)
- Kaggle AI Agent Securityコンペ振り返り ー345th Place Solution imp 0 / dev 75 / cod 75 / har 70(いいね相当スコア: 4)
- フレーム問題とは?AIの古典的難問 imp 0 / dev 60 / cod 30 / har 30(いいね相当スコア: 0)
- OCEL 2.0 を使って実業務でプロセスマイニングするときの3つの課題 imp 0 / dev 70 / cod 65 / har 30(いいね相当スコア: 0)
- numpyでAttentionを実装したら、√dで割るのは万能ではなかった imp 0 / dev 85 / cod 80 / har 25(いいね相当スコア: 0)
- Kaggle Pokémon TCG AI Battle Challenge ポケカコンペ振り返りーメダルなし imp 0 / dev 60 / cod 60 / har 20(いいね相当スコア: 9)
- Instagramのリーチ予測モデル -「新規アカウント収集・モデル再学習」の自動化と監視を実装した imp 0 / dev 80 / cod 80 / har 50(いいね相当スコア: 1)
- 25億token学習したGPT-2 Smallに、さらに25億token学習させたら性能は上がるのか? imp 0 / dev 80 / cod 75 / har 20(いいね相当スコア: 0)
- YOLO26nアーキテクチャ徹底解析 step1 | 全体像 imp 0 / dev 85 / cod 80 / har 15(いいね相当スコア: 1)
- 音声モデルを量産して分かった18のこと imp 0 / dev 85 / cod 85 / har 50(いいね相当スコア: 取得失敗)
- Agentic RAGの自律判断を深掘り:LangGraphの思考プロセスとツール連携 imp 0 / dev 85 / cod 85 / har 80(いいね相当スコア: 0)
- Ollamaの使い方マニュアル imp 0 / dev 75 / cod 65 / har 40(いいね相当スコア: 0)
- 「人間のほうがAIっぽい」を検算したら、記事が4倍長いだけでした imp 0 / dev 40 / cod 35 / har 25(いいね相当スコア: 0)
- 画像生成AIのしくみ(第一話) imp 0 / dev 75 / cod 50 / har 20(いいね相当スコア: 1)
- 化合物の水溶解度を機械学習で予測してみる③:LightGBM で精度向上 imp 0 / dev 75 / cod 75 / har 15(いいね相当スコア: 0)
- AIはなぜ質問に答えられるのか?知っておくべき3つのメカニズム imp 0 / dev 70 / cod 40 / har 30(いいね相当スコア: 0)
- AIに自動で意識研究をやらせて見えてきたこと imp 0 / dev 70 / cod 70 / har 75(いいね相当スコア: 取得失敗)
- 長期AIエージェントの履歴を一度だけ構造化——監視精度0.48→0.85〜0.87、120段課題は8/30→30/30 imp 0 / dev 85 / cod 85 / har 85(いいね相当スコア: 取得失敗)
- AIはもう「賢さ」を競う段階ではないのかもしれない――TRON-AIという日本の空席 imp 0 / dev 60 / cod 50 / har 50(いいね相当スコア: 取得失敗)
- 第二話「ジェミちゃん登場なのだ!」 imp 0 / dev 30 / cod 15 / har 20(いいね相当スコア: 取得失敗)
- Xで見かけたLLM関連の話題を時系列で並べてみた(8月中旬〜9月頭) imp 0 / dev 55 / cod 45 / har 40(いいね相当スコア: 取得失敗)
- ローカルLLMのJSON出力が空で返る — 壊れ率0%を100%に戻した設定は format:"json" ではなかった imp 0 / dev 80 / cod 80 / har 65(いいね相当スコア: 取得失敗)
- RTX 5090+128GB RAMとDGX Spark(GB10)×2でQwen3.8 Flash-Nextを動かして比較してみた imp 0 / dev 75 / cod 60 / har 40(いいね相当スコア: 取得失敗)
- カリフォルニア2026年会期のプライバシー・AI立法 / 積み上げ型規制の到達点とAI監査人登録制 雑感 imp 0 / dev 40 / cod 25 / har 30(いいね相当スコア: 取得失敗)
- 人間は合理的な正解だけでものを選ばない――AIに人間味が必要な理由 imp 0 / dev 70 / cod 60 / har 75(いいね相当スコア: 取得失敗)
- StartLux-V1.0-27B-Previewとは? imp 0 / dev 70 / cod 60 / har 50(いいね相当スコア: 取得失敗)
- ペット保険の先行実験場 / 予防の内部化とAI全自動査定 雑感 imp 0 / dev 50 / cod 50 / har 40(いいね相当スコア: 取得失敗)
- Fable 5からFable 5.1へ乗り換える前に見る3点。料金は上がらず、枠は増えない imp 0 / dev 80 / cod 75 / har 65(いいね相当スコア: 取得失敗)
- 4世代を通して見る Gemini Flash — 3.5 から 3.8 までの4か月で何が変わったか imp 0 / dev 80 / cod 70 / har 50(いいね相当スコア: 取得失敗)
- #9 🌙ルナと学ぶ、本当に初めてのローカルLLM|オフラインAIで最新情報を検索?LM StudioにDuckDuckGoを追加する imp 0 / dev 70 / cod 65 / har 50(いいね相当スコア: 取得失敗)
- 【論文】精神科AIを診断と対話で測る imp 0 / dev 75 / cod 65 / har 30(いいね相当スコア: 取得失敗)
- ローカルAIとLLMの現在地、創作の裏側を覗く夜|AIとわたしの深夜エンタメ会議 #3 imp 0 / dev 50 / cod 35 / har 30(いいね相当スコア: 取得失敗)
- AIコンパニオンとの恋愛のライフサイクル / 事業者が握る別れと法 雑感 imp 0 / dev 40 / cod 30 / har 35(いいね相当スコア: 取得失敗)
- 【生成AIニュース+】『Gemini 3.8 Flash』『ChatGPT Images 2.1リーク』『H3 Max Turbo(preview)』『H3-World』『Muse Spark 1.3』『OpenVDN / VDN-H3』『ComfyUI-VDN-H3』『Video Delta Net』『MiniMax H3向けVR180 SBS LoRA』『MiniMax FastH3 on Reactor』『H3 Motion Context 0.5.0』他多数 imp 0 / dev 70 / cod 50 / har 40(いいね相当スコア: 取得失敗)
- 撮られる側の同意 / スマートグラスと監視の日常化 雑感 imp 0 / dev 30 / cod 20 / har 25(いいね相当スコア: 取得失敗)
- Qwen3.5 9B Q4 VS Qwen3.8 27B Q2|RTX 3060 12GBで実際に比べてみた imp 0 / dev 80 / cod 70 / har 50(いいね相当スコア: 取得失敗)
- 黒パグ再起動第一弾(続くとは言ってない)LLM48の世界 バグアイドルが世界を席巻するまで(大げさ) imp 0 / dev 35 / cod 25 / har 25(いいね相当スコア: 取得失敗)
- 記事自動生成パイプラインが壊れた 5 箇所──プロンプト起因は 2 件だった imp 0 / dev 80 / cod 85 / har 75(いいね相当スコア: 取得失敗)
- Java誕生から現代に至る歴史をゴスリン氏ら関係者が語る公式ドキュメンタリー「The Java Story」、YouTubeで公開中 imp 35 / dev 20 / cod 20 / har 5
- Webブラウザ上でターミナルのシミュレータを実行、GitやCLI、Vimなどを無料で学べる「WebTerm Learn」が公開 imp 40 / dev 65 / cod 60 / har 30
- AWS、新たな太平洋海底ケーブル「Sta'O'Nuk」発表。420Tbps、2029年に稼働へ。集中リスク排除のため新たな陸揚げ拠点を採用 imp 50 / dev 20 / cod 10 / har 5
- AWSをゲームで学べる「AWS Cloud Quest」に新バージョン「AWS Cloud Quest 2.0」登場! AIによるバーチャル顧客と対話し、要件を聞き出して正しくソリューションに落とし込め imp 35 / dev 60 / cod 50 / har 20
- AWSとAzureが最大100Gbpsでの相互接続を開始。これでAWSはAzure、Google Cloud、Oracle Cloudとのマルチクラウドをサポート imp 45 / dev 50 / cod 40 / har 20
- ボットの振る舞いを動的に学習して防御を改善し続ける、「Adaptive Intelligence」ボット検出エンジン、Cloudflareが発表 imp 55 / dev 75 / cod 75 / har 60
- ClaudeにSalesforceを統合した「Claudeforce」、AnthropicとSalesforceが発表。Claudeから営業データ分析や顧客対応を実現 imp 60 / dev 80 / cod 80 / har 75
- iOSDC Japan 2026にLINEヤフーのエンジニア7名が登壇します(9/11〜9/13) imp 20 / dev 30 / cod 20 / har 0
- MIRU2026参加レポート imp 40 / dev 70 / cod 65 / har 25