AI News Digest 2026-09-08
対象期間: 直近2日間 / 収集記事 718 件
台本で使った記事
特集
- Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
- Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- Train What You Deploy:Token-Faithful Post-Training of a Production Coding
開発者コーナー
中堅コーナー
AIツール紹介コーナー
速報コーナー
- AI Cyberattack Capabilities Reach 'Critical' Level… OpenAI and Anthropic Restrict Access, While South Korea Moves at a 'Snail's Pace' - finance.biggo.com
- OpenAI reports AI "research interns" and warns about its own pace at the same time
- Presentation: From AI Agent Demo to Production: Automated Testing and Evaluation
- Our security team is suddenly very interested in prompt injection now that agents can take actions
- A 6-hour successful agent task isn’t really 6 hours of autonomy
- Which security checks are actually missing from current AI agent APIs?
- I tested Pydantic AI for 3 days on saas.pet's content QA agent — here's the honest take
- Grok 4.5 Cuts Coding Costs 80%, Trails Rivals [2026] - tech-insider.org
- How Figma Uses AI Agents for Security
参考記事一覧
参考記事一覧を表示
- Huawei Prepares 160,000 Ascend 950DT Accelerators for DeepSeek Data Center - TechPowerUp
- North Korea-linked hackers turn to AI agents for malicious email decoys - Korea JoongAng Daily
- Despite xAI’s objections, judge rules not to block Minnesota “nudification” law - WDIO.com
- Research acceleration: The view inside OpenAI — NVDA Quantitative Valuation Record
- Supporting independent journalism in Ukraine — NVDA Quantitative Valuation Record
- The Dataflow Model Revisited
- Grok 4.5 Cuts Coding Costs 80%, Trails Rivals [2026] - tech-insider.org
- Grok simulates the final outcome of the Denver Broncos’ 2026 season ahead of Week 1 against the Kansas City Chiefs - A to Z Sports
- bzip3
- Watch Los Angeles get built, one building at a time (1880–2026)
- Keep Our Servers Running
- AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
- Caltech Mathathon – first hackathon ever devoted to research level mathematics
- PostgreSQL 19 Interactive Tour
- Live map of public transport in Belgium
- Speculative Decoding in vLLM on AMD GPUs
- Smartphone makers don't bother to comply with EU repairability requirements
- Impedance Matching (2017)
- Volkswagen to convert German car plant to produce Israeli defense equipment
- Splash-free urinals (2025)
- Tell HN: OpenAI brings back 5 hour limit for plus and business standard users
- Making a Python interpreter in 1024 bytes
- Ask HN: How do you manage skills files?
- Bill Gates tries to install MovieMaker
- GrapheneOS Overhauled Default Apps and Secure Clipboard
- Apparently CodePen 2.0 sends data to their servers as you type
- LG smart TVs caught logging audio with screen off and snooping on local devices
- An Alien Mind
- Kotlin 2.4.20 Released
- Java Annotated Monthly – September 2026
- The Rider 2026.3 Early Access Program Is Open
- The complex corporate web behind a $3.2 billion AI data center
- Authors push back as publishers and agents make claims on Anthropic settlement
- Travis Kalanick’s Atoms might be getting into the robotaxi business
- Seattle Times and Newsday sue OpenAI and Microsoft for infringement
- Presentation: From AI Agent Demo to Production: Automated Testing and Evaluation
- Google Mantis: an Agentic Vulnerability Scanning Harness for Reducing False Positives
- How Figma Uses AI Agents for Security
- Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- GPT-6 Astra beat Portal start to finish without human help in under 24 hours
- ChatGPT claws back web traffic share to 55.5 percent as Gemini's brief comeback fades
- AI-designed drug appears to turn back the body's biological clock in early trial
- How AI wiped out an entire industry in Nairobi
- At UBS, AI skills are now a condition for landing a job
- OpenAI reports AI "research interns" and warns about its own pace at the same time
- New York City bans AI tools from public schools through eighth grade
- Qwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver
- Chatbots built an "echo chamber of one" and now psychiatry has to decide if "AI psychosis" exists
- Last Week in AI #343 - GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1
- EXAONE Forecast for Finance
- From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
- Data-Optimized Contingency Screening: A Machine Learning Approach to Power System Security
- Iris: Climbing to the Search Frontier
- A Removal Based Approach to Improve LLM Faithfulness at Test-Time
- Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
- Corporate Language Model (CLM): Transforming Tacit and Fragmented Enterprise Knowledge into a Sovereign, Auditable, and Executable Corporate Intelligence Layer
- HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
- When Quantization Breaks Memory: Recurrent-State Write-Back in Low-Precision Temporal Inference
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality
- Rethinking Indirect Prompt Injection as a Test-Time Search Problem
- BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents
- MaxKernel: Agentic Kernel Generation for TPUs
- Towards a universal language of concepts: A survey
- Data-Driven Discovery of Composition-Dependent Constitutive Models for Hyperelasticity and Viscoelasticity of Digital Materials
- From Answers to Interpretations: Rethinking Ambiguity-Induced Aleatoric Uncertainty Estimation in LLMs
- IPGeoAI: Transformer-Based Geolocation with LLM Semantic Fusion
- Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
- La Agente \'Optima: Towards Agentic Self-Driving Laboratories
- Extremely Sparse Supervision Incentivizes Reasoning Ability
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines
- $\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
- Leveraging Imperfect Restoration for Data Availability Attack
- SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents
- A Cost-Aware Agentic Architecture for NL-to-SQL over Nested Enterprise Schemas, with a New Benchmark
- Continual Graph Memory for Adaptive Recommendation under Intent Drift
- Harness-agnostic detection and immunization of reward hacking in self-evolving language models
- ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
- Train What You Deploy:Token-Faithful Post-Training of a Production Coding
- Predicting Spatiotemporal Mobile Sensing-Based PM2.5 Concentrations Using Low-Rank Adapted Spatially Attentive Graph Neural Network
- SQL-Zero: Self-Evolving Text-to-SQL
- Model Retirement Creates Reproducibility Risk in Biomedical AI Publications
- FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
- PLUME: Parameter-Efficient Personalization of Large Language Models via Low-Rank User Modulation in Shared Subspaces
- Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM
- DCFA: Dual-view Causal-inspired Attribution for Failure Reasoning in LLM-based Multi-agent Systems
- Shadow Queries for Private Retrieval in Vector Databases
- Diffusion Language Models for Mobile Edge Agentic AI: Foundations, Applications, and Challenges
- DODR: Deterministic Operator-Driven Reasoning in Latent Space
- ProtLingo: Efficient Protein Language Modeling via Conditional Memory and Expert Routing
- Whose record is this? Diagnosing and authorizing record use in personalized multimodal models
- Hierarchical Possession-Aware Graph Pointer Network for Pass Receiver Selection
- MedFlow: Class-Aware Multi-Scale Generation for Medical Time-Series Synthesis
- When Financial Fine-tuning Fails: A Three-Level Detectability Analysis of Numerical Hallucination in Domain-Adapted Language Models
- CPR-IE:A Compression-Prediction-Resource Intelligence Efficiency Metric
- Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance
- ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults
- MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models
- MZ-Rain: Moisture-Budget-Guided Zero-Inflated Model for Station-Level Precipitation Nowcasting
- CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
- LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents
- CHAMP: Cross-domain Hybrid Architecture for Matchmaking and Prediction in Online Multi-Player Games
- AutoLR: Automating the Path from Research to Launch Review in Industrial Recommender Systems
- MARLA: A Conceptual Scaffold for Regulatory Learning under the EU AI Act
- Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach
- From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments
- Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
- Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing
- Solving Hard XAI Queries Based on a Compiled Dual-Rail Encoding
- Why We Care About Understanding: Competence through Predictive Compression
- Global to Local: Topology-Preserving Adaptive Graph Pooling via Granular-Ball
- A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering
- Language models judge war differently when tested for alignment
- TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing
- Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
- Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications
- MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning
- TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
- Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
- Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent
- LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
- ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding
- Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation
- Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
- SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
- A Hybrid Predictive Ensemble of Machine Learning and Deep Neural Networks for Early Cardiovascular Disease Risk Assessment
- The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior
- What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection
- CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review
- ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
- Substrate-Aware AI Agents: Execution Context as a First-Class Input
- Uncensored Open-weight Models: Redistribution as the Persistence Layer
- Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory
- A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability
- Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
- Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents
- AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
- Testing Interchangeability in LLM Agent Teams
- GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity
- Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
- RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
- Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness
- LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams
- Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
- Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability
- Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
- Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
- Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
- Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
- A Deep Generative Model for Synthesizing Labeled Wireless Signals
- AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks
- When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs
- Scalable Context Orchestration for Serving LLMs Over Voice
- Evidence Integration in Large Language Models
- Abstraction Agent
- Data-Driven Learning of Unknown Nonlinear Differential Equations Using Functional Analysis
- Adapting from Downturns: Prediction of Long-Term Conversational-Skill Development in Mental-Health Crisis Counselors
- VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models
- Blockchain-Enabled Secure Logging for Fiscal Electronic Mechanisms: Evaluation of the Greek eSEND and myDATA Tax Systems
- Cross-modal triage network: a multimodal deep learning framework for severity-based triage and visual explainability in chest radiographs
- Ultrasound-Based Prediction of Cirrhosis Decompensation Using Large-Scale Computer Vision Models
- Where Appearance Fails, Geometry Recognizes: A CAD-Free 3D Shape Prior That Complements Vision Foundation Models
- What Moves? Localized Motion Representations for Compositional Scene Control
- You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
- A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
- REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation
- GRACE: Graph-Grounded Reflective Agent Copilot Engine for Expert-in-the-Loop Knowledge Expansion
- When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models
- A Roadmap for MEG Foundation Models
- Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
- Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation
- Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
- Towards Understanding Pause Token Fine-Tuning Dynamics: A Mode Retention Perspective
- Hakken: Predicting future discoveries to fill the gaps in today's knowledge
- A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap
- Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters
- Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes
- Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection
- Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI
- Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics
- Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models
- When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models
- Dual-Part Multi-Lateral Branched Network for Multi-Class Segmentation in Cardiovascular Catheterization Angiograms
- PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning
- SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
- Tracing Audio Grounding and Answer Selection in Audio LLMs
- Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle
- Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based Fusion
- Wireless Foundation Models: State-of-the-Art and Open Challenges
- Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty
- Building a research-software catalog with a coding agent: from hackathon prototype to public deployment
- Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models
- Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
- When Does an Interpretation Count as Established? The Formation, Evaluation, and Responsibility of Interpretation in Generative AI
- Persistent Teacher Anchoring for Tool-Using Agents
- Dynamic Heterogeneous Graph Representation Learning: A Survey
- Can Activation Steering Capture Multidimensional Authorship Style?
- Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents
- Recurrence Is Not Enough: Causally Validating Multilingual SAE Translation Features in Gemma 2 and 3
- Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution
- Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
- MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate
- MMTClinic: Multimodal, Multilingual Time Series Question Answering and Reasoning Benchmark for Clinical Domain
- CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
- Mitigating Performance Discrepancy in Cross-Domain 3D Class-Incremental Learning
- PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation
- Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents
- ReCAST: Restoration-aware Cascaded Stage-wise Training for Obfuscated SMS Risk Classification
- SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection
- Attention-guided super-resolution of 4D flow MRI in carotid arteries
- RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
- Adaptation Interfaces for In-Context Tabular Foundation Models in Time-to-Event Prediction
- Sound-based Multi-Person 3D Pose Estimation
- Methane Detection On Board Satellites from Unorthorectified Imagery
- Better Understanding, Better Fixes? A Study of Hallucination in LLM-based Automated Program Repair
- TreeFI: Value-Aware Statistical Fault Injection for Deep Neural Networks
- ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems
- One Diffusion Model, Two Roles: Guided Trajectory Planning and Safety-Critical Scenario Generation in Closed-Loop Simulation
- MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
- VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
- Amortizing Scaling Law Construction Costs
- How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI
- Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection
- How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions
- Qlippy: A Retrieval-Augmented GenAI Assistant for Reproducible Quantum Workflows and Experiment Tracking
- Beyond Co-purchase Relation: Evolution of Complementary Recommendations at Allegro
- Adaptive Multi-Granularity Temporal Modeling for Weakly Supervised Video Anomaly Detection
- A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support
- NEAT-POCKET: Pocket-Conditioned Autoregressive 3D Molecular Generation with a Neighborhood-Guided Set Transformer
- TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors
- A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
- Beyond Stationarity in Time Series: Discovering Causal Structures and Latent Regimes via Markov Blankets
- AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics
- Phase Transition Frequency as a Training Time Predictor of Test Accuracy in ResNets
- A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR
- PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting
- Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
- CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls
- How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing
- RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
- Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions
- The History Is the Detector: Executing CVE Patch History, End-to-End
- Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
- When LLM Decompilers Recompile More and Preserve Less
- What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies
- Reflection-aware Generative Novel View Synthesis
- RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments
- Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction
- Quality-diversity in dissimilarity spaces
- A Survey on Semantic Modeling for Building Energy Management
- Procedural Content Generation via Generative Artificial Intelligence
- Active Inference for an Intelligent Agent in Autonomous Reconnaissance Missions
- Achieving Olympiad-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
- RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training
- MemCoRe: Recovering Evidence from Progressively Compressed Factual Knowledge for Agent Memory
- OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Heuristic Design
- The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
- MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution
- BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA
- Role-Aware Artificial Intelligence Across Augmentation and Automation in Human-Machine Symbiosis
- Robust and Efficient Guardrails with Latent Reasoning
- SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
- Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents
- GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
- EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
- Improving Weak World Models Behind Strong Agents in Atari Pong
- Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
- KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
- NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
- $A^2E$ : An End-to-End Agent Auditing Engine
- Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
- AI Revealed Preferences
- FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
- SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning
- Agentic Context Cracking: Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
- Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades
- Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
- LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
- Measuring proximity to standard planes during fetal brain ultrasound scanning
- An Empirical Study into Clustering of Unseen Datasets with Self-Supervised Encoders
- Hyperedge Anomaly Detection with Hypergraph Neural Network
- Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
- AI-Powered CPS-Enabled Vulnerable-User-Aware Urban Transportation Digital Twin: Methods and Applications
- Graph Foundation Models for Recommendation: A Comprehensive Survey
- AccidentSim: Generating Vehicle Collision Videos with Physically Realistic Collision Trajectories from Real-World Accident Reports
- Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
- Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
- Gradient-based Model Shortcut Detection for Time Series Classification
- Partial Inverse Design of High-Performance Concrete Using Cooperative Neural Networks for Constraint-Aware Mix Generation
- GLOW: Graph-Language Co-Encoding for Agentic Workflow Performance Prediction
- The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI
- TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
- Multi-Modal Time Series Prediction via Mixture of Modulated Experts
- Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace Adjustments
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Post Fusion Bird's Eye View Feature Stabilization for Robust Multimodal 3D Detection
- MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
- SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
- YOLO with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection with trustworthy multimodal AI in computer vision perception
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
- CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
- Octopus Protocol: One-Shot Hardware Discovery and Control for AI Agents via Infrastructure-as-Prompts
- HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
- Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
- Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases
- Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs
- FVSpec: Real-World Property-Based Tests as Lean Challenges
- "**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems
- From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- SoK: AI-Augmented Binary Reversing
- Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization
- Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
- Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
- Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Deep Divide-and-Reduce in Symbolic Regression
- TRNet: Learning with Topographic Priors for VHR Paddy Rice Mapping
- Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
- LEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNN
- Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction
- Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity
- Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms
- MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters
- DeMMO: Longitudinal and Cross-Disease Modelling of Digital Mobility Outcomes via Multi-Task Learning
- When Linguistic and Internal Confidence Diverge in Large Language Models
- E-SENS: Exclusion-Sensitive Penalization for Negative-Constraint Retrieval
- SPD: Single Pass Decoding for Generative Reranking
- DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models
- From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
- Post-Training Language Models for Gold-Medal Performance in Coding Competitions
- Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings
- Air-Ground Collaborative Vision-and-Language Navigation via Shared Bird's-Eye Maps
- IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks
- Almost Free State Prediction Separation
- FWBC-VLA: Force-Aware Whole-Body Compensation for Contact-Rich Loco-Manipulation
- Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR
- Spectral-Target Physical Latent Structuring for JEPA-Style World Models
- ProToMEx: Rapid, Interpretable Explanations via Structured Representations
- A Data Fusion Framework for Grounding Aerospace Surrogate Model via Experimental Wind-Tunnel Observations
- Quantum-Assisted Memory-Efficient Training for Parameter-Intensive Wi-Fi-Based Human Activity Recognition
- Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning
- BER-PEF: Unified Human Mobility Predictability Evaluation via Bayes Error Rate Estimation
- Modular Deep Recurrent Neural Network: Application to Quadrotors
- SharedSAE: One Feature Dictionary Across Language Models
- A Quantum Variational Approach to Prototypical Recurrent Unit
- On the Abundance of Critical Points of the t-SNE Energy
- Disentangling Attention in Deep Operator Learning: A Controlled Study of Data-Driven and Physics-Informed Architectures
- Beyond a Universal Forecasting Selector: Demand-Conditioned Model Selection across Demand Patterns and Horizons
- A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
- Conformity Breaks Conformal Prediction
- On-board ML for Trace Gas detection in Imaging Spectroscopy data
- Nested Inductive Bias Framework for SPD Manifold Learning
- An Energy-Based Conservative-Dissipative Latent Neural Evolution Operator for Magnetization Dynamics
- Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
- Mitra-v2 Technical Report
- Fast Surrogate Modeling of Excitable and Oscillatory FitzHugh-Nagumo Dynamics with Parametric Neural Operators
- Optimizer Memory Schedules for Outscaling the Overtraining Axis
- Representation Redundancy and Structural Complexity in Finite-Field Inversion
- GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer
- Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator
- SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery
- Interpretability for Turing Machines
- WEECFP-SuRGE: Wide Embedded Extended Connectivity Fingerprint with Substructure Rotary Graph-distance Encoding
- Locating and Steering Refusal Beyond Attention
- Training Large Language Models for Small-Molecule Design with Synthetic Task Scaling
- A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision
- Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning
- A Robust Watermark-based Fingerprint Framework for GNNs Ownership Verification
- Learning-Augmented Algorithms: Guarantees, Construction Mechanisms, and System-Level Implications
- How Faithful Is Attribution for Sales Forecasting? A Counterfactual Study
- Federated Attack Campaign Detection via Contrastive Encoding of Threat Indicators in Gradient Updates
- Communication-Efficient Personalized Federated Learning via Layer-Wise Multi-Threshold Random Sketching
- PACE: Propagation-Aware Collaborative Correction for One-Shot Personalized Federated Graph Learning
- KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU
- When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation
- From Deep to Shallow: Unconstrained and Efficient Layer Merging Strategy
- Fast Gauss Sums via Flash Attention
- Physics-Aware Random Walk Fingerprints for Scalable Power Grid Graph Classification
- Fractal basins trap latent reasoning
- BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
- Beyond Homoscedasticity: Decoupled Uncertainty Optimization for Deep Imbalanced Regression
- Solution-space heterogeneity shapes federated learning dynamics across partial differential equations
- Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks
- Deep Microcompression: Structured Pruning and Bit-packed Quantization for Microcontrollers
- A Comparative Study of Counterfactual Explainers for Graph Neural Networks Enabling Multiple Types of Graph Edit
- Single-Query Black-Box Calibration Auditing via Logit Bias
- Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy
- MomentQuant: an even more minimalist interval method with linear time complexity for time series classification
- From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline
- Dimension-Adaptive Batched Lipschitz Narrowing Without Knowing the Zooming Dimension
- FedDRAW: Federated Dual Reputation Annealing Weighting for Heterogeneous Multi-Institutional Chest Radiograph Classification
- Hessian-based molecular conformation augmentation for a scalable and efficient strategy of machine learning interatomic potentials
- GLASS: Graph-Language Alignment with Spherical Scoring for Transferable Graph-Level Anomaly Detection
- How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method
- Learning from VAE Errors to support ECG-based Differential Diagnosis of Myocardial Scar
- Optimal Rates for Agentic Networked Information Aggregation
- Embedded Graph Flows for Categorical Graph Generation
- Variational Continuation for Double Pendulum Periodic Orbits
- Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
- Interface-Induced Trajectory Censoring
- GEPARD - Generative, Prosody-aware, Autoregressive text-to-speech model for Realtime Dialogue
- Self-Supervised Pretraining of Molecular Graph Encoders with LeJEPA
- Low-Latency Spell Correction for Japanese Music Search Queries
- Compute-in-Memory Attention: A Time-Domain Analog Softmax Circuit with RC-Tunable Temperature
- Corporate-Family Resolution Is Not a String-Matching Problem: A Public Benchmark Stratified by Name Visibility
- TNFlow: Amortized Posterior Inference for Trans-Neptunian Object Surface Composition
- The microscope is the mask: privileged views and labels from a cryo-ET forward model
- A Constraint-Aware Generative Framework for Synthetic Origin-Destination Demand in Logistics Networks
- Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys
- Candidate Comparability Before Promotion: Conditional Validation in Adaptive Network Intrusion Detection
- Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters
- Recovering molecules from coarse-grained beads: free-energy-conditioned generative backmapping across chemical space
- Client-Side Probing of Deleted Ridge Statistics in Federated Unlearning
- Scale-QLoRA: Code-Invariant Adapter Merging for Native 4-bit Microscaling LLMs
- A Sim-to-Real Study of Surface-Code Decoder Benchmarking
- MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning
- Centered Permutation Prefixes for SGD with Random Reshuffling: Sharp Rates, H\"older Geometry, and Composite Proximal Extensions
- Hidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR
- Latent-Aligned Reasoning for Multimodal Recommendation
- A Differentiable Neural Surrogate for Photon Propagation in Neutrino Telescopes
- LookThere! Sparse Vision by Reinforced Selection
- Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware
- Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving
- Minimax Lower Bound for Estimating Diffusion-based Local Intrinsic Dimension
- Coupled Control and Wireless World Models for Resilient Remote Robotic Control
- An Analysis of Self-supervised Pre-training with Dependent Samples
- Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks
- Conformal Prediction for Offensive Security
- SMILE: Self-Explainable Multimodal Information Bottleneck for Medical Diagnosis
- FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search
- PAC-Bayesian Reconstruction Guarantees for Time Series Variational Autoencoders
- Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC
- Shallow neural network approximation in mixed Sobolev spaces
- LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics
- Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning
- Adaptive Gated Deepfake Detection for Low-Resolution and Resource-Constrained Environments
- UniMate: One Unified Model to Animate Diverse Skeletons
- Small Molecule Optimization with Large Language Models
- The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures
- Explainable Clustering of Mixture Models
- DeltaGNN: Graph Neural Network with Information Flow Control
- TSMini: A Simple Yet Highly Effective Trajectory Similarity Learning Model
- Towards Efficient Parametric State Estimation in Circulating Fuel Reactors with Shallow Recurrent Decoder Networks
- Deep Learning-Driven Peptide Classification in Biological Nanopores
- WaveletDiff: Multilevel Wavelet Diffusion For Time Series Generation
- Fractal and Chaotic Activation Functions in Echo State Networks: Preprocessing Topology Governs the Echo State Property
- Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks
- Consensus Group Relative Policy Optimization for Text Generation
- Brain4FMs: A Benchmark of Foundation Models for Electrical Brain Signal
- Reservoir-Based Graph Convolutional Networks
- The Geometry of Polynomial Group Convolutional Neural Networks
- Advancing Subseasonal Forecasting with Machine Learning
- Relocation of compact sets in $\mathbb{R}^n$ by diffeomorphisms and linear separability of datasets in $\mathbb{R}^n$
- Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications
- Inductive Venn-Abers and related regressors
- Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
- Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models
- Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective
- From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing
- Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale
- Latent Fact-Checking: Detecting Misinformation through Activation Engineering
- Boosting Data Augmentation with Stochastic Weight Averaging
- ClosureBench: A Constructive Benchmark for Compositional Graph Reasoning
- Across-Design Uncertainty in Short Pricing Panels: Inference and Identification
- Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
- Canalization Before Generalization: Grokking as a Dynamical Probe
- TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision
- Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions
- OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items
- TrajMind: Chaining Role-Specialized LoRAs for Fast-and-Slow Collective Trajectory Anomaly Diagnosis
- A Location-Invariant Estimator of Extremal Quantile Treatment Effects for Heavy-Tailed Distributions
- GraphMend: Code Transformations for Fixing Graph Breaks in PyTorch 2
- Constrained Sensing and Reliable State Estimation with Shallow Recurrent Decoders on a TRIGA Mark II Reactor
- Enhancing Affine Maximizer Auctions with Correlation-Aware Payment
- Regularity of Second-Order Elliptic PDEs in Spectral Barron Spaces
- Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics
- Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control
- Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
- Second-order consistency for learning chaotic dynamics via randomized Jacobian matching
- Quantum Kolmogorov--Arnold representation theorem for continuous unitary-valued maps
- Statevector-to-Hardware Reconstruction of a Four-Qubit ZZ Quantum Kernel: A Single-Backend Case Study of Three Execution Jobs
- Automatic knot selection in smooth additive models
- To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
- Token-Level Advertising
- An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis
- Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
- Omega-N: Interpretable Structural Node Descriptors and Their Applicability Domain
- SocialBuddy: Tailoring Search Agent for Social Scenarios
- Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
- DTM: Deterministic Approaches for Black-box Test Suite Minimization with Tree-based Similarity
- Breaking the Alphabet: Rethinking File Ordering in Code Review
- AI Writes Code, Humans Pay the Debt. An Empirical Study on the Sustainability and Evolution of Agent-Generated Code
- The Prompt Triangle: A Registered Report on Prompts as Hybrid Artifacts
- SH-PDOPS: AI-Driven Cloud Native Enterprise Reliability Framework for Predictive Analytics and Intelligent DevOps Automation
- Data-Related Challenges and Requirements for Event Log Generation in Process Mining: A Systematic Literature Review
- Big Questions on Software Architecture: Report of the ICSE 2026 BoF on Software Architecture
- Toward Model-Driven Digital Twin Configuration: Separating Structure Semantics and Runtime with SysML SAREF and Ditto
- A Mixed-Method Empirical Study of LLM Assistance in Software Engineering Workflows
- A systematic literature review on logging smell detection
- Engineering as Code: Bringing Software Engineering Methodology to Engineering Design
- A Governance Methodology Layer for AI-Assisted Software Development: Defect Taxonomy, Controlled Ablation, and Process-Over-Capability Evidence
- Large Language Models for Fuzz Testing in Microservices: A Systematic Literature Review
- Robustness and Trade-offs for Code LLMs on Protected Code
- Automated Deployment of Real-Time Tasks for Phased Execution on Scratchpad-Based Multicore Platforms
- Physics-Direct FPGA Tooth-Contact Computation for Deterministic Gear Digital Twins
- Reviewer Capability Governs Rejection Targeting, Not Repair Skill: Evidence from LLM Execute-Review-Revise Pipelines
- Integrating Crash Report Mining and LLMs for Bug Localization and Repair: An Industrial Report
- An Empirical Analysis of CodeQL False Positives and Query Refinements for Java Vulnerabilities
- Software Engineering in the Agent Era From Trustworthy Change to Human Agent Software Organizations
- How Developers Discuss Generative AI: A Longitudinal Study of the Visual Studio Code Community
- An Empirical Study on Learning Paths and Gender Dynamics in Scrum Master Roles
- T(r)opical Islands: Visualizing & Understanding Socio-Technical Artifacts
- Ritgard: T(r)opical Islands of Socio-Technical Artifacts on GitHub
- CPL: A Compact C-like Systems Language with Explicit Low-Level Control
- Adaptation Needs in Robotic Systems: Assessing Behavior Trees and Their Enhancement
- Guiding AI to Fix Its Own Flaws: An Empirical Study on LLM-Driven Secure Code Generation
- Evaluating Uncertainty and Quality of Vision-Language-Action-enabled Robots
- EmoPyLab: A Tensor-Native, Hardware-Accelerated Laboratory for High-Throughput Benchmarking and Decision-Making in Multi/Many-Objective Optimization
- The Unseen Delta: Characterizing the Compiler Optimization Landscape via Top-Down Differential Analysis
- The Neglected Baseline in Model Interpretation
- A faster way to convert a timestamp to Hour, Min, Sec
- The shortest IPv6 addresses
- It took a year to ship WebAssembly in Anubis
- Two Tiny Utils for the Result Pattern
- Internationalization and Localization
- Python Iceberg
- Rust debugging survey 2026 results
- How well do agents use test/verification techniques?
- Demystifying complex configurations
- QBittorrent breaks out of sandbox to commit crimes
- "Hammock Driven Development" (2010)
- GEM for Linux provides a classic graphical desktop with windows, menus, dialogs & a 68K emulator
- Following legal advice, the Nitter project will continue
- What are you doing this week?
- NetBSD 11 from scratch
- What every kernel programmer should know about Jump Labels
- Have the frontier labs mixed up AI safety and security?
- Terence Tao on “prematurely solving [a maths] problem by purely AI-powered methods”
- "Simple Made Easy" (2011)
- Data races and the limits of ThreadSanitizer in C and Go
- Debian Code Search: Fast TurboPFor with Go SIMD
- Protecting Mission-Critical Data Beyond The SoC: Why Inline Memory Encryption… — NVDA Impact Analysis & Price Prediction
- Intelligent Engineering: From Optimization To AI — NVDA Impact Analysis & Price Prediction
- El desarrollo de software en la era de la IA: de escribir sintaxis a diseñar arquitectura
- BizNode Pro: run up to 5 independent Telegram bots, each with its own identity, knowledge base, and AI persona
- Chip Industry Week In Review — NVDA Impact Analysis & Price Prediction
- Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod — NVDA Impact Analysis & Price Prediction
- A Practical Desk Workflow for Free AI Image Generator Drafts in 2026
- The Complete Guide to Agent-to-Agent Marketplaces in 2026
- ‘NBA 2K27’ With NVIDIA DLSS 5 Leads 26 New Games Coming to GeForce NOW — NVDA Quantitative Valuation Record
- I tested Pydantic AI for 3 days on saas.pet's content QA agent — here's the honest take
- The model did the reverse-engineering. The validator was the hard part.
- Your system prompt isn't instructions. It's data.
- Changes to LLM pricing: Baidu and Tencent
- Your Gaming PC Just Became an AI Powerhouse: Meet FreeToken!
- Build a Local LLM Chatbot with Ollama and Python
- Ollama Guide: Eigene KI lokal hosten statt Cloud-Abo – So geht's
- Uncensored AI chat online: text without a moderation layer, and the honest limits
- An LM Studio alternative for Mac that needs nothing installed
- Run Qwen 3.6 without a GPU: the hosted route and what it costs
- Use Hermes 3 405B online, in the browser or from any OpenAI-compatible client
- Top 463 Sites to Buying Edu Emails for Student Discounts 2019
- I REALLY hope the new gemma 5 family sticks to the "chat model first" philsophy and doesn't fall into the Qwen trap
- MiniCPM5-2B Release Day
- After over a year of my nights and weekends, the Jenny app is done!
- Are you running Qwen 3.8 27b or Qwen Flash Next?
- 9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled
- DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds!
- New Benchmark: The Struggle Bench
- Why are the SOTA open-weight models scoring (relatively) low scores on AA-Omniscience Index
- How to squeeze out every last drop of your precious RAM on your Mac - Use iPhone mirroring
- when will open source LLM catch up to Astra I wonder?
- Higher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark
- tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval)
- The models are fine, our toolings and methods are shit.
- little-coder vs just Pi
- Qwen 3.8 Next Flash is really really REALLY verbose..
- Cybersecurity is local AI model's killer use case
- Benchmarking calories evaluation with LLMs
- Which models are you running on 32Gb VRAM (16+16) and 128Gb RAM?
- Is anyone running dual GPUs with x570 or x870 Taichi?
- Qwen Next on 24 + 64 GB VRAM?
- Local LLMs and their use case on your specific hardware
- Radeon RX 7900 GRE 16GB and RX 480 8GB Vulkan benchmarks llama.cpp
- Easy local Copilot with VS Code and Lemonade
- 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
- Weekly Hiring Thread
- Just got laid off from Agentic AI Firm, So giving away my 3 years worth of LinkedIn content marketing playbook for Agentic AI for free
- Here's a free tool to stop AI from forgetting anything !!
- I gave an AI agent $50 and 24 hours to book meeting leads. It ended up roasting 40 founders, getting a 60% reply rate, and making $600.
- If you were starting from zero with AI today, what would you actually learn first?
- A 6-hour successful agent task isn’t really 6 hours of autonomy
- Should agent frameworks define your agent? I’d love feedback on A11
- Nvidia CEO says "AGI has arrived" after GPT-6 Astra. Are we actually there, or are we moving the AGI goalpost again?
- Astra Smokes Fable 5.1!!
- Did Claude/Codex actually replace no-code tools, or am I missing the point?
- What "production-ready" actually means, explained for founders who don't code
- (Genuine Question) At what point does weird AI agent behavior become an actual security incident?
- What are the most important concepts to know about AI for a software developer?
- No idea if a script is going to cost me 20 cents or 20 dollars until it's finished
- Open-weights LLMs vs frontier APIs: when to rent, when to own
- Agents keep hitting the same wall: missing boring tools that are easy to plug in
- What is one AI task you would trust completely, and one you would never trust AI with?
- I want to take this to the next level
- Our security team is suddenly very interested in prompt injection now that agents can take actions
- Best free/open-source resources for building automated agents locally (24GB RAM + RTX 2060S)?
- Ling 3.0 (on Nous)
- Running AI agents on WhatsApp taught me why human-in-the-loop isn't optional. The platform itself punishes you for autonomy
- Sistema automatizado de trade com metatrader 5
- Which security checks are actually missing from current AI agent APIs?
- The purpose of DNS is to spread scams
- There's No Limit to How Bad Code Can Get
- Quoting Zach Kehs
- Remind
- Clipnote
- Airuncode
- Tucky
- Assist
- An Interview with Zhipu AI Co-founder Tang Jie by Qiushi Magazine - Geopolitechs
- Tough choices ahead as the US pushes a global artificial-intelligence divide - Bruegel
- Chen Danian’s AI Model Nearly Surpasses DeepSeek Within 3 Months Post Launch - eu.36kr.com
- Every Major AI Model That Shipped in Six Weeks, From Kimi K3 to GPT-6 Astra - Memeburn
- China's Development Model: Building a Sustainable Global Foundation for Shared Progress - eu.36kr.com
- No Longer Follow, But Define: The "DeepSeek Moment" of China's Household Ventilators - eu.36kr.com
- China's Open-Source AI Models Become the Global Backbone: Saudi Arabia, Japan, and Europe 'Wrap' Chinese Weights - finance.biggo.com
- AI Cyberattack Capabilities Reach 'Critical' Level… OpenAI and Anthropic Restrict Access, While South Korea Moves at a 'Snail's Pace' - finance.biggo.com
- The AI threat has eclipsed sci-fi - eKathimerini.com
- OpenAI rolls out GPT-6 Astra - TahawulTech.com
- Year-End Bitcoin Price Bets Get Wild as 11 AI Models Target up to $105K - Cryptonews.net
- Canva Turns AI Chatbots into Growth Engines - varindia.com
- How To Use SpaceXAI's Grok Build - Engadget
- Tesla Optimus, Grok, and TSLA Stock: Inside the Physical-AI Bet That Now Defines Tesla - InvestorPlace
- DW News. . Elon Musk's AI chatbot Grok is facing lawsuits over alleged AI-generated child sexual abuse material. Now, a new lawsuit makes an even more serious claim. #dwdigital - facebook.com
- Musk-Backed xAI Plans Doosan Power Plant - Businesskorea
- ChatGPT, Claude and Grok Stopped Working. Here’s Why - Memeburn
- GROK $0.0{5}8536 | Live GROK Price Chart Today, Swap on USDT | MEXC - mexc.co
- Elon Musk Grok AI Predicts a Possible Solana Supercycle by 2027 - 99Bitcoins
- RTX 5090でllm-jp-4-33b-thinkingをEXL3量子化・評価した2日間の振り返り
- Langfuse v3 から v4 への移行で、既存のトレースと親子関係はどう変わるのか
- 【SALT2 エンジニアリング勉強会】ワークフローからAI Agentへ:仕事を任せるための境界と検証の設計
- AIに却下済みの設計を3回蒸し返されたので、「前例」を配るMCPサーバーを作った
- 2026年9月版:主要LLM 料金・スペック・共通ベンチマーク比較
- 【2026年版】生成AIを学ぶための51冊 ― LLM・RAG・AIエージェント・LLMOps・Physical AIまで
- 自然言語仕様をコンパイルする──Compile by Trainingで再利用可能な神経関数を生成する
- Pythonで作るAIエージェントのTools入門
- Agentic Loopをゼロから組んでみる
- AI彼女アプリを作っていて気付いた。どうやら時代が追いついてきた
- Astraの正直さを測るテストに、空欄の答案を出してみた
- クレカ不要の無料LLM APIを実運用する — レート制限の読み方とフォールバック設計
- ローカルAIに仕事を回す仕組みが、3週間動いていなかった
- Geminiをアップデートしたら、日本語PDFの品名だけが消えた
- Prompt Steering Replacement: プロンプトによる介入をSteering Vectorで置き換える
- 欲求は知能をどう駆動するか——進化が作った人間の動機と、AI時代の新しい適応
- AIは、そこにある関数をもう一度書く — RepoExecが測った「依存を再利用する力」
- 量子化LLMのtokens/sだけでは見えないCTIRとE2E latency
- RAGの心臓部 — Embeddingとベクトル検索で「意味」を捉える
- ログ肥大・トークン制限・パース失敗 ― ループを回し続けると必ず出会う3つの壁
- Peyman Milanfar「A Lagrangian View of Flow Matching」を読む
- Jaffray Woodriffから学ぶSystematic Trading
- メーターの数字は、写真だけでは読めない
- 誤差逆伝播を「触って」理解する — 数字が動くまで、式は腑に落ちない
- TimesFM 3をColabで試す:関連データや予定を足すと予測は変わる?
- YOLO26nアーキテクチャ徹底解析 step2 | Conv, DWConv, PWConv ブロック
- 🐔ナナフシからロボットの歩き方を学ぶ
- BigQuery ML × 時系列モデルARIMA_PLUSでEC売上の週次予測を自動化する
- GPT-2-likeからQwen2-likeへの実験:第3回 絶対位置EmbeddingをRoPEに変える
- FXのトレンド発生は予測できる?EURUSD・15分足を全足検証
- 音声AIで「考えるほど無言になる」を防ぐ:推論コストをターン別に振り分けるTypeScript実装
- ClaudeがLeanでフェルマーの最終定理を11日間の自律実行で完全形式化証明
- LLMの思考能力向上:推論時コンピュートの仕組みと活用戦略
- 【備忘録】Hermes AgentでGPT-5.6系をCodex OAuth経由で利用する際のProvider指定メモ
- ARIMAモデルをJavaScriptでスクラッチ実装した
- 才能にたよらない耳コピ -- 8年ぶりのAI対決
- dera AI Weekly Vol.47 — 2026/9/7
- AIというラベルでは責任は決まらない / 中国最高人民法院のAI紛争裁判規則 雑感
- 生成AIと児童保護法制 / 規範を先に書く者 / 自主的枠組みと測定可能性 雑感
- 通信量はAI英会話でどれくらいか。音声処理が端末とサーバーのどちらで走っているかで決まる
- 第五話「疲れている人に5000文字の励ましを送ってはいけない」
- レッドラインは何を禁じるのか / AI自己申告と独立検証 / 2026年という期限 雑感
- AIの安全性と法 / 監視可能性の逓減 / 安全基準を誰が引くのか 雑感
- 海外におけるAI保険動向 / 助言から行為へ / AI免責条項の拡大と保険募集該当性 雑感
- 自動運転と法 / 介入データから交通規則を学ぶ自動運転 / 継続学習と時点固定型規制の距離 雑感
- もっともらしさの出どころ / ウェアラブル健康推論ベンチマーク 雑感
- 情報の規制から関係の規律へ / 米国12州コンパニオンボット法 雑感
- トルコのAIガバナンス / 調達要件としての基本的人権影響評価 / トルコ2026-2030年AI行動計画へのCAIDP意見
- 推論モデル時代のエージェント設計:テスト時計算量とツールの融合
- 2026-09-06 Hacker News Top 10
- 【生成AIニュース+】『OrcaRouter Qwen3.8-27B-Uncensored』『Lightning v3.1 Pro』『Shot Composer』『ComfyUI-Ref2VA-VSA』『NaughtyTimes-MiniMax-H3』『ComfyUI MiniMax H3 NegPiP』『Poppy』『Iris』『Blender Cascadeur RT』『screenwriting-skills』『Qwen3.8-Flash-Next-DS4-IQ2』『Hypershell Halo』
- nanochat#8 GPT-2の点数は 0.256525 だった
- 今さら聞けないAI用語|「生成AI」と「LLM」は何が違う?
- Gemini3.8Flashってすごくないか?
- もう他人事じゃない!AGIと自立型AIがもたらす社会のパラダイムシフト
- 教えていないのに生えたもの(Opus4.6)
- 毎日エラーゼロのAI自動化パイプラインが実は週5日しか動いていなかった
- 【パンダ船長の技術航海-7号】(令和8年08月17日)観測は、失敗を誰のせいだと記録するか――原因の帰属と出所の証明
- ああ、女神(AI)さま
- 【GPT-6 Astra解読】AIが「思考を隠し始めた」――史上初のサイバーCritical到達と開発現場が知るべき多層防御
- 語彙力も表現力もこれ一つ!GPT-6 Astraでリサーチの質を爆上げする方法
- コーナーだけ減速?Fusionの送り最適化を使い直す|Fusion CAM
- インターンシップ選考に落ちたら、その会社はもう受からない?~26卒データサイエンティストがインターンシップ選考の見送りから、本選考で内定をつかむまで~
- 巨大オンラインゲーム「EVE Online」。ゲームを支える240万行のPython 2.7のコードをPython 3へ移行すると発表
- HTMX 4.0正式リリース。内部実装がXHRからfetchに移行しStreaming HTMLが可能に、属性はデフォルトで子要素に継承されないように変更など
- VS Code誕生から現在までの物語「The Story of VS Code」YouTubeで公開。作者のエリック・ガンマ氏はなぜIBMからMSへ移籍してVS Codeを作ることになったか
- JavaScriptをC言語にコンパイルする「porffor」がアルファ版に到達。ネイティブやWASMを生成可能
- 社長の考えをいつでも聞ける!AIコジーを社内公開