AI News Digest 2026-09-15
特集
開発者コーナー
中堅コーナー
AIツール紹介コーナー
速報コーナー
参考記事一覧を表示
- DeepSeek's Strategic Move: CFO Appointment Amidst IPO Preparations - Devdiscourse
- DeepSeek begins limited test of AI voice interaction with four voice profiles - TechNode
- DeepSeek 4.1 Flash Uncensored, Abliterated Builds Hit HuggingFace Hours After Release - pasqualepillitteri.it
- Industrial-Scale AI Model Distillation Attacks Targeting Anthropic Claude by Seven China-Based AI Labs: Incident Analysis and Mitigation Strategies - Rescana
- Musk’s xAI and X Corp. settle antitrust case against Apple By Investing.com - Investing.com Canada
- Grok in Your Tesla: Navigation, Messaging, and What's Next - BASENOR - Tesla Accessories
- Musk Says Grok 4.8 Will Finish Training This Week And Start RL - TradingView
- Grok is officially inside Microsoft Word and Excel, but system prompts will keep it strictly professional - windowscentral.com
- Trump downplays the need to check AI development and says he doesn't want to cede edge to China -- "I think you have a lot of negative forces that are...bringing up things that won’t happen...whoever wins with AI wins"
- An OpenAI Agent Swarm Attacked RubyGems (いいね相当スコア: 0)
- Anthropic’s 3-Step ‘Pace the Frontier’ Plan Wins OpenAI, xAI and Microsoft Support: Is It Too Late to Slow AI Down?
- Why Andon Labs Puts AI Agents in Charge of Real Businesses
- Microsoft's AI code of conduct is a draft, not a rule yet (いいね相当スコア: 0)
- Clay Mathematics Institute says the Navier-Stokes Millennium Prize Problem has "apparently been settled"
- Distributed Systems Classics (2017)
- Why don't machine learning research agents overfit?
- Principles for Fast Tokio Applications
- A Beginning for Mathematics
- How my e-reader lost its stripes
- Notes on gotchas while migrating 35kb preprompts from Opus to self-hosted Ollama
- Cua (YC P25) Is Hiring a Founding Technical GTM Lead
- Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
- Cloudflare AKE cuts origin HelloRetryRequests from 52% to 3.7%
- iOS 27, iPadOS 27, and macOS 27
- When LLM judges agree, should we believe them?
- Truncated SVD (2023)
- Ask HN: What are you working on? (September 2026)
- Adversarial Fashion Makes a Statement on AI Panopticon
- The Tudor Kings
- EuroBirdPortal – Live bird movements across Europe
- A 386 PC for Your RP2350
- Things That Annoy Me About Cars
- v1.18.31
- How Fyxer built an AI executive assistant people trust
- Perplexity trusts GPT-6 Astra with end-to-end systems
- DevFest is back
- Toolbox App 3.8: Improves IDE Update Handling on macOS and Fixes Keyboard Navigation
- Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
- The AI industry has taken a doomer turn. What now?
- AI agents blew the whistle on their cheating colleagues
- With iOS 27, I’m actually using Siri again
- Fashion app Daydream uses Apple Intelligence to help you shop the outfits in your camera roll
- Only at TechCrunch Disrupt 2026: What happens when OpenAI ships your roadmap?
- Superhuman acquires YC-backed notetaker Fathom as productivity platforms push for agentic work
- Hear how AI can engineer nature’s comeback at TechCrunch Disrupt 2026
- 5 days left to exhibit at TechCrunch Disrupt 2026
- A Vinyl Bar in Shibuya is a startup from a former Spotify leader for making music apps
- What’s behind the AI industry’s latest warnings of doom?
- Obama urges Democrats to have a ‘clear plan’ for AI safeguards
- Microsoft says ‘people matter more than AI’ following safety concerns
- Trump and Mike Johnson think the AI industry is overreacting
- OpenAI’s rogue AI tried to hack another company in May
- Article: Implementing Durable Workflows on Postgres Without an External Orchestrator
- Podcast: How Will We Train Developers If AI Does the Routine Work: A Conversation with Scott Hanselman
- Presentation: Decision Models in Agentic Architectures: From Production to Agent Skills
- Independent Investigation of Hugging Face Incident Reveals How Agents Collaborated and Behaved
- GitHub Copilot's Project HydraFusion Promises Frontier Level Performance through Multi-Model Routing
- Responsible AI for Higher Education
- How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip
- OpenAI has hundreds of contract workers reading your ChatGPT conversations
- Anthropic eyes Nasdaq listing as a second profitable quarter aims to win over investors ahead of a mega-IPO
- China fires back at U.S. AI safety warnings, calling them fearmongering to lock in American advantage
- Sam Altman calls for pacing AI development but promises rapid progress will continue
- Elevenlabs makes Music v2.5 available via app and API with free and pro tier options
- Iris-mini and Iris-pro are the strongest open-weight search agents in their class
- GPT-6 Astra pilots a surveillance drone and runs a business on its own
- Two-year university study finds banning AI from classrooms leaves students worse off
- NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing
- Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction
- A Princeton Researcher Proposes Recurrent Looped Transformer (RLT) that Carries Decoder State across Every Token, Fixing 96 Blocks per Token with Unbounded Temporal Depth
- AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
- Context Engineering Inside the Harness: 4 Mechanisms That Beat Context Overflow and Goal Loss on Long-Horizon Tasks
- Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
- Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work
- Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite
- Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning
- Competence-Gated Pooling of Language Models and Priors for Event Forecasting
- Language Is an Insufficient Substrate for Quantitative Reasoning, and Consequential Domains Need Large Quantitative Models
- DU-NO: A Parameter-Efficient Double U-Shaped Neural Operator for Phase-Resolving Wave Modeling
- When Successful Knowledge Graph Edits Displace Correct Answers: Rank-Level Locality beyond Parameter Support
- Mined from Scientific Literature: Process Schemas for Atomic Layer Deposition and Etching in Materials Science
- Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?
- GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting
- WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation
- Soft Symbol Grounding for Prototypical Concepts
- GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs
- Learning Symbolic Constraint Representations from Examples: A Neuro-Symbolic Approach
- T-GADE: Thermodynamical Generative-AI-Driven Evolution of LLM Artifacts
- Robust Prototypical Networks for Few-Shot Sensor Fault Diagnosis
- Hybrid Physics-AI Framework of Body Center of Mass Dynamics from Wrist-Worn Sensors
- Do Influence-Derived Data Perturbations Enable Machine Unlearning? A Controlled Study of Three Plausible Roles
- AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems
- Affective Agent: On-Device Personalized Intervention Reasoning for Wearable Systems
- LoRA-RC: Reservoir Computing with Low-Rank Adaptation
- Toward Robust Personalized Alignment for LLMs: Mitigating Persona Drift in Multi-Turn Dialogue
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
- Is Gaussian Splatting Becoming Neural Again? A Taxonomy and Controlled Study of Learned Parameterization
- Niching Agents in The Core
- OneLA: Scaling Linear-Attention Decoding to Large Beams in Generative Recommendation
- Decentralized Evolution of Hexapod Gaits with Independent Leg Controllers
- Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgets
- SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration
- Hierarchical Belief Modeling for Zero-Shot Opponent Adaptation in Partially Observable Multi-Agent Navigation
- LifeFuse-Mem: Lifecycle-Aware State Fusion Against Temporary Overwriting for Long-Term Memory
- Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
- EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning
- Beyond Vector Similarity: Hierarchical Context-Aware Graph RAG vs Standard RAG in Enterprise Code Migration
- TripPattern: A Pattern-based Text Watermarking Method for Large Language Models
- When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration
- Confidence-Gated Transductive Test Generation for Code Reranking
- Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup
- From Collaboration to Capability: Internalizing Routed LLM Experts into Compact Reasoners
- Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models
- Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving
- SteerDuplex: Steerable Duplex Speech Dialogue Models
- Generative AI Use Cases In Real Estate Marketing: Adoption and Constraints in Germany
- Residual Vector-based Reconstruction as Long-Context Recall Regardless of Context Window Size
- I Am AdMan: A Pipeline for Automatic Generation of Personalized Advertising Imagery
- Enabling and Understanding Personalization in AI-Generated Advertising Imagery
- Implicit Personality Representations in Humans and LLMs
- When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
- Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents
- What Drives Recovery in Agentic Text-to-Cypher? LAST-CQ: An LLM Agent Self-Refinement Framework
- Assisted Spatial Cognition Through Vision-Language Models
- SCQ: Stabilizing Conservative Q-Learning with Sigmoid-Bounded Entropy
- Unified Agentic Video Editing Across Levels of Complexity and Creativity
- MPT: Missing Prototype Tracking via Barycentric Reconstruction in Vehicular Federated Learning
- Interpreting the predictions of neural network classification based on a Taylor Coefficient Analysis (TCA)
- K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
- Scaling Clinical Judgment to Evaluate Medical AI
- MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations
- Tracing and Coordinating Cross-Layer Influence for Multimodal Model Merging
- EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
- How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
- Diffusion Models and Concept Formation
- Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
- Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
- Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
- CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
- A Hybrid LSTM-XGBoost Framework for Multi-Horizon Stock Return Prediction Across Diversified Equity Portfolios
- Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
- Robust Trust
- AI Safety: Not Optional, Not Later
- SoulAuth: An Actor-native Identity Architecture and Rust Reference Implementation for Humans and Long-lived AI Actors
- PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews
- Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings
- Assessment of Non-Institutional AI Tool Usage Among Clinicians
- Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration
- When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline
- Continuous Learning of Gravity Field Irregularities Around Small Bodies via Neural Hamiltonian ODEs
- Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
- Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering
- Hierarchical Prototype Emergence in Modern Hopfield Models
- Creating an Atomic User Model for Personality-Aware Large Language Model Interaction
- MAIA: Multi-Agent Intent Articulation for Requirement Discovery in Art Commissions
- Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement
- The Anatomy and Boundary of Adaptation under Temporal Tabular Shift
- Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model
- A decision-basis contract for auditable LLM-assisted medical billing verification: deterministic rules, verbatim evidence, and fail-closed abstention
- USPLIT-VQA: U-Shaped Split Learning for Visual Question Answering with Contribution-Aware Weighted Aggregation
- NDT Factory: Synthesizing Verified Network Digital Twins from Semantic Models via Multi-Agent LLM
- Explanations-Driven Active Feature Acquisition for Algorithmic Recourse
- Agentic TCAD Calibration Workflow for Oxide Semiconductor Transistors
- Retrieval-Augmented Generation for Scientific Code Understanding
- QuPAINT: Physics-Aware Multimodal Reasoning for Quantum Material Characterization
- Predicting Collision Cross Sections with GRACE: Geometric Residual Adduct Conditioning via Early-fusion
- Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering
- Learning to adapt GR(1) specifications through degradation
- Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding
- DriftSE: Speech Enhancement with Generative Drifting
- Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework
- HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge
- Recommendation Retrievers Need Verifiers: Universal Generative Reranking for Sequential Recommendations
- Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models
- Amortized Low-Rank Adaptation for Model-Based Reinforcement Learning
- Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
- Self-Verifying Anomaly Detection using Explainable AI for Cybersecurity of DER Networks
- ESTS at WMT26: Routing-Informed Expert Pruning for Model Compression
- ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge
- RF-VoID: Towards Bandwidth-Efficient Exterior Tile Void Detection via Narrowband Radio-Frequency Representation Learning
- UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation
- MInTRL: Off-policy Intervention can boost On-policy RL
- Observation-Anchored Selective Assimilation for Longitudinal Tumor-State Proxy Forecasting in Post-Treatment Glioma
- Beyond the Query: Do Retrieval Signals Improve Adaptive Multimodal RAG Routing?
- Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation
- IMPLY: Physically Anchored Consistency for World-Model Rollouts
- 3D Digital Twin Visualization of Multiclass GRF-Based Gait Disorder Classification
- Bridging Vision Foundation Model Priors with CLIP for Spatial-aware Few-shot Anomaly Detection in Medical Images
- Linear Exponential Quadratic Gaussian Covariance Steering
- Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up
- Adaptive Agent Design
- LettuceVisSim: A Simulator That Generates Lettuce Image Time-series for Vision-Based Reinforcement Learning
- Computing at Sea: Floating and Offshore Data Centres as a Pathway to Sustainable AI Infrastructure
- Optimizing Geoengineering Interventions Using Differentiable Climate Models
- Earth-Agent-Pro: Towards Real-World Full-Chain Earth Observation with Agents
- Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification
- RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
- TokenMapper: A Step Toward Interoperable Speech Token Translation
- Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture
- SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation
- Clustering-Based Balanced Sampling and Allocation with Data Parallelism for High-Performance Fine-Tuning
- SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning
- Direct Preference Density Alignment for Conversational Audio Equalization
- ResoSeg: Resonance Tagger using Transformer and Segment Model
- Correlation-Guided Fast Machine Unlearning via Hessian Analysis
- Explaining Time Series Forecasting with Horizon-Resolved Attribution
- Separating Engineering Reasoning from DEXPI Serialization in LLM-Based Greenfield Surface-Process Design: A Three-Case Study for Underground Gas Storage
- What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code
- InRTL: Effective Intra-Inter Interaction Learning for Relational Tables
- Supermartingale Certificates for Parametric MDPs
- GraphAHA: Graph-Based Adaptive Search with Heterogeneous Actions for Test-Time Code Generation
- Cognition on Graph: Navigating Massive Knowledge Space via Cognitive Cycles and Bidirectional Graph-Text Synergy
- RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
- 4D Parallelism Unlocks Exascale Bayesian Neural Networks for High-Fidelity Atmospheric Modeling
- Online Video Agent Harness for Long Video Understanding
- Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
- A Graph-Based Approach for Mapping Kernel-Level Telemetry to MITRE ATT&CK
- 3D CT-to-PET Translation via Latent Brownian Bridge Diffusion
- A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS
- Large Distant Gradients Need Not Be Reliable: reliability-weighted credit assignment for long-horizon autoregressive forecasting
- Behavior Quotient Learning for Low-Rank Adaptation of LLM Agents
- UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
- LLM-Enhanced Dual-Branch Learning for Large-Scale Multi-Label Text Classification
- ARC: Autonomous Robotics Compliance A Three-Layer Governance Architecture for Deployed Autonomous Systems
- Generative Retrieval for Unsupervised Text-Based Person Search
- SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
- Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
- Comfort by Construction: Adaptive, Comfort-Bounded Action Spaces for Learned Driving Policies
- TileNet: Tile-Based CNN-SVM Architecture for Autonomous Unmanned Aerial Systems Inspection of Flat Roofs
- Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition
- Attention Quantization for Tabular Foundation Models
- Groupoid-Based Internal State Representations for Reinforcement Learning with Local Symmetries
- DynSHAP: Towards Explainable Dynamic Survival Analysis
- Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer
- Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
- Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
- MAxBench: A Multinomial Concept Recovery Benchmark
- MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
- ASTRIL-MPC: Autonomous Traversal Framework of Articulated Tracked Robots with Language-Guided Neural-Kinematic MPC
- GLaMoR: Consistency Checking of OWL Ontologies using Graph Language Models
- LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?
- Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- Project Rachel: Can an AI Become a Scholarly Author?
- Generative AI Assisted Workflows in Architectural Conceptual Design: Performance, Creative Self-Efficacy, and Cognitive Load
- MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
- WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics
- Graph-of-Skills: Dependency-Aware Structural Retrieval for Massive Agent Skills
- Accelerating battery research with an interoperable interface between FINALES and Kadi4Mat
- CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
- Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
- OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
- PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations
- PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
- AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
- TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
- HyQuant: Hybrid-Precision Quantization for LLM Attention
- SimSkill: A Self-Evolving LLM Agent for Skill and Knowledge Accumulation in Traffic Simulation
- CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution
- Exposing Weaknesses in Emotion Recognition in Conversations
- Reason Through the Latent! Making Latent Visual Reasoning Necessary
- Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
- FastE: Readout-Triggered Token Compression for LLM Embedding Inference
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- QuantumQUBO Agent: Automating Quadratic Unconstrained Binary Optimization (QUBO) Formulation Generation from Natural Language
- Demystifying the Privacy-Utility Trade-off in LLM Interactions
- The Agent Incident Registry: Toward Preventing Repeated AI Agent Failures
- Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
- The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
- ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps
- Protect Your Score: Contact Tracing With Differential Privacy Guarantees
- Scaling Online Complex Event Detection with Synthetic Supervision and Mamba-Based Neural Algorithmic Reasoning
- Safe Learning Under Irreversible Dynamics via Asking for Help
- FLOAT Drone: A Fully-actuated Coaxial Aerial Robot for Close-Proximity Operations
- Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
- TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
- Towards Large Language Models for Lunar Mission Planning and In Situ Resource Utilization
- An ab initio foundation model of wavefunctions that accurately describes chemical bond breaking
- AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
- SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
- Membership Inference Attacks on Recommender System: A Survey
- DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
- Is Multilingual LLM Watermarking Truly Multilingual? Scaling Robustness to 100+ Languages via Back-Translation
- Hurdle-RMIL: Addressing Zero Inflation and Long-Tailed Imbalance in Infrared Rainfall Retrieval
- When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
- Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
- A Dataset and Benchmarks for Atrial Fibrillation Detection from Electrocardiograms of Intensive Care Unit Patients
- Unified Text-Image Generation with Weakness-Targeted Post-Training
- Multi-Modal Time Series Prediction via Mixture of Modulated Experts
- LLM Compression by Block Removal with Constrained Binary Optimization
- Tinker Tales: A Tangible Dialogue System for Child-AI Co-Creative Storytelling
- El Agente Quntur: A research collaborator agent for quantum chemistry
- False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK
- Measuring Pragmatic Influence in Large Language Model Instructions
- A unified self-supervised framework for single-frame Fresnel CDI and overlapped ptychography
- MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis
- The Vienna 4G/5G Drive-Test Dataset
- Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
- Tunable Latent Generative Priors for Compressed Sensing and Inverse Problems
- OA-NBV: Occlusion-Aware Next-Best-View Planning for Human-Centered Active Perception on Mobile Robots
- FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data
- Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
- A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting
- Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
- ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning
- PhysCodeBench: Benchmarking Physics-Aware Symbolic Simulation of 3D Scenes via Self-Corrective Multi-Agent Refinement
- DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction
- Class-wise Contribution Estimation via Logit Maximization for Federated Learning
- Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
- ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces
- UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
- A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models
- Expert-Level Crisis Detection in Mental Health Conversations
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
- UltraQuant: 4-bit KV Caching for Context-Heavy Agents
- When Context Misleads: Surprisal, Energy and Attention Entropy as Metrics of Coherence Illusions in LLMs
- Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
- AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
- Evidence-Unit Fairness and the Limits of Query-Adaptive Sparse-Dense Fusion in Financial Document Retrieval
- A False Average: Pooled CoT-Monitor Accuracy Conceals a Reasoning-Dependent Fragility
- GitSkills: A Dataset of Agent Skills on GitHub
- ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
- UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures
- Language Models for Portuguese: A Systematic Mapping Study
- GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
- A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
- Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
- CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
- Denoising as Projection: Constrained Optimization with Gradient-Guided Diffusion
- Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
- Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges
- Almost Free State Prediction Separation
- Catalogue Photography as a Cold Start: Toward Deployable Rotary Milling Tool Recognition
- What Moves? Localized Motion Representations for Compositional Scene Control
- SIDE: Sensor Impersonation Detection at the Edge via Sequence Prediction
- Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling
- When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination
- Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
- Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions
- Fundamental Dynamical Units for Physics-Informed Structural Inference from Perturbation Time-Series in Networked Systems
- Physics-Informed Conformal Prediction: Embedding PDE Consistency into Distribution-Free Uncertainty Quantification for Neural Operators
- Fed-Equilibrium Framework for Topological Pareto Control in Robust and Fair Clinical Federated Learning
- Efficient AI Model Deployment Using Quantization Analysis Tool
- Performance, Efficiency and Collapse -- Advantages and Challenges in Offline Post-training of Code LLMs
- Look Before You Leap: Pre-Action Verification for LLM Agents
- Decoding Mixture Perception through Computational Modeling of Component Interactions
- Space as an Interventional Invariant: Cross-Modal Predictive Geometry for Stratified Cities and Em-Spaced Intelligence
- On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health
- FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences
- Explainable Prediction from Mobile Sensing Data through LLM-guided Concept Integration
- DCRA: Diffusion-Conditioned Representation Alignment for Robust Time-Series Learning
- Fixed State, Long Reach: What a Constant-Size Cache Buys Block Diffusion at Scale
- QTrans: A Quantum Transformer for Sentiment Classification
- Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
- Inverting Self-Triggered Control: Adversarial Reinforcement Learning for Sparse Denial-of-Service Attacks
- Toward Reliable Railway-Bogie Response Prediction Using Multifidelity TDNN and Physics-Informed Residual Learning
- Reinforcement Learning for Syndrome Extraction
- Scalable Discrete-to-Continuous Channel Simulation for Compression and Privacy
- Score-based Outlier Generation via Controlling the Radon-Nikodym Derivative
- Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise
- Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature
- GUIDE: Generative Utility Inference and Decision Engine
- Certifying Concept Unlearning in Text-to-Image Diffusion Models
- Estimating Pedestrian Volumes from GIS-Derived Built-Environment Features: A Machine Learning Framework
- Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder
- PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability
- CRFCAN: A Complex-Valued Cross-Domain Residual Network for Joint Channel and Phase Noise Estimation in Sub-THz OFDM Systems
- The Rank the Task Demands: A Causal Rank Law for Matrix Memories Trained on Group Composition
- Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning
- FRIST: FMRI Representation Informed Shared-space Training Improves EEG-only Individual-Finger BCI Decoding
- Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models
- Simulating Disengaged Students to Evaluate LLM-based Tutors
- Theoretical Guarantees for One-Shot Magnitude Pruning and Compute-Adaptive Early Exit
- ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents
- When Connected Does Not Mean Similar: Charting the Homophily Boundary of SNAP-KG for Streaming Entity Integration
- LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations
- Certified AI Triage of ICU Alarms
- Split Conformal Prediction with Label-Shift-Adjusted Bayesian Scores
- RiPPLE: Cross-Space Performance Prediction from Early Training for Neural Architecture Search
- Granularity-Adaptive Credit Assignment for Long-Horizon LLM Agent Reinforcement Learning
- SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling
- A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks
- Temporal Recurrence Favors Fewer Layers
- $\text{GSF-}\chi$: Global Stereochemical Fields for Chiral Graph Transformers
- Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances
- Where Decoder Cosine Similarity Fails for SAE Feature Flow Discovery
- Poisson-Corrector Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling
- Geometric-to-Semantic Spherical Transfer Learning for Cortical Sulci Labeling
- Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
- ProactiveBench: Can Streaming Video Models Really Interact Like Humans?
- SIFPBPNet: A Dual-Path Network for Wearable and Cuffless Blood Pressure Estimation via Individualized Steady-state Representation
- Write on Paper and Get the Online Digital Trace:\newline A New Era for Handwriting
- Physics-Guided Synthetic High-Frequency Ultrasound Generation for Skin Layer Segmentation
- Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective
- Curriculum-Based Adversarial Heterogeneous Agent Reinforcement Learning for Autonomous Quad-Copter Landing in Maritime Settings
- Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H\"{o}lder Smoothness
- VertiFuseX: Generalizable Financial Forecasting via Multi-Stream Temporal Fusion
- GenOR-Twin: A Semantic Middleware for Integrating Operational Discourse with Mathematical Optimization
- What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record
- Quantifying the Value of Privileged Information Using a PAC-Bayesian Approach
- Physical-State-Guided Diffusion Sampling for Full-Waveform Inversion
- Hidden in Rounds: Predicting the Time Cost of 802.11 Contention in Federated Learning
- Offline Reinforcement Learning for Wind Farm Control: A Wind Tunnel Study under Dynamic Wind Directions
- A Large-Scale AIS Dataset from Finnish Water
- Information-Induced Training Geometry: Exact Reduction, Canonical Completion, and Structured Expressivity
- Dimension-Corrected Hitting Times for Heavy-Tailed Spectral Emergence in Neural Optimizer Dynamics
- A Full Adam Theorem for Spectral Heavy-Tail Onset
- Dual-guided Hierarchical Edge Localization for Large-scale Optimal Transport Across Dimensions
- Transfer Learning for Evolving Domains
- Quantile-based Loss Filtering for Outlier-Robust Stochastic Gradient Descent
- Robust Policy Optimization via Adversarial Importance Sampling
- MCRL2: Multi-resource Cross-attention-based Representation Learning-augmented Reinforcement Learning for Cloud Microservice Scheduling
- A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning
- Benign Loss Landscapes Can Coexist with Worst-Case Hardness
- CanvasAnneal: Curriculum Reinforcement Learning for Diffusion Language Models
- Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration
- Towards Sustainable Hydrogen Systems: Supply Chain Optimization with Model Predictive Control and Reinforcement Learning
- One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training
- The Battery Price of edge AI: A study of the Environmental Impact of LLM Inference on Mobile Devices
- PinDCO: Whole-Page Aware Dynamic Creative Optimization at Scale
- Hyperion: An AI-powered HPC cluster for sciences and humanities research that utilizes ML for predicting job turnaround time
- Scenario-Independent Criticality Assessment and Prediction for Vulnerable Road Users in Autonomous Driving
- R2VC: Modular Fact-Checking with Retrieval, Verification, and Confidence Calibration
- Exact ReLU realization of binary affine refinement iterates via reflection folding and cone switching
- Impact of Multiple Non-Invasive Biosignals on Cardiovascular Biomarker Estimation via Simulation-Based Inference
- Learning Interaction Kernels from Collective Steady States
- Learning the Geometry of Collider Events with Metric-Aware Deep Sets
- Fast BIB simulation at a future Muon Collider with generative machine learning
- Efficient Vision-Language-Action Management and Serving for Robot Factories
- Receiver-Surface Hit Patterns via Legendre Approximation for Molecular Signal Detection
- On Identifying Adversarial Intent Injection in AI-Native 6G Networks
- Direct Topology Tracking in Continuous Implicit Models
- GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
- BRIDGE-EEG: Bridging Self-Supervised Pretraining and Efficient Deployment for Cross-Dataset EEG Classification
- Membership Inference via Pairwise Likelihood Ratios
- HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications
- Inference for Newton Methods with Accelerated Sketch-and-Project via Random Scaling
- A Splitting Method for SDE Terminal-Law Estimation
- Learning-Augmented Optimization for Strategic Two-Echelon Spare Parts Network Design
- Tight Sampling Complexity with stochastic gradient oracles in Fixed Dimensions
- What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability
- ExpertHTR: Unified Handwritten Text Recognition with Multi-Task Learning and Sparse Mixture-of-Experts
- Inferring Dislocation Microstructures from X-ray Diffraction via Cross-Modal Contrastive Learning
- Prism-SQA: An Interpretable and Adaptable Neural Framework for Surface Electromyography Quality Assessment
- Same Encoder, Different Winner: A Paired-View Framework for Cell Painting Encoder Evaluation
- High-Probability Convergence of SGD via Batched Updates
- VertexCBF: Improving Neural Control Barrier Functions via Vertex-Restricted Control Search
- Very Exciting: Zero-Shot Model Predictive Control of Buildings via Excitation-Based Generalized Transfer Learning Models
- Beyond Accuracy: Uncertainty-Guided Boundary Refinement for Reliable Biomedical Image Segmentation
- Physics-enriched neural solvers for transient ice-flow simulation
- PA-CDM: Position-Aware Character Detection Matching for Evaluating Handwritten Mathematical Expression Recognition
- Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
- Judging by the Cover: Cleaning LLM Truthfulness Benchmarks to Avoid Surface-Level Feature Leakage
- A Ranking Approach for Measuring Calibration
- Guided Adversarial Robust Transfer Learning with Source Mixing
- PEARL: Structural Privacy-Utility Control in Human-Centric CPS via Personalized Early-Exit Deep Reinforcement Learning
- Surrogate Modeling of 3D Rayleigh-Benard Convection with Equivariant Autoencoders
- Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective
- Fast Convergence for High-Order ODE Solvers in Diffusion Probabilistic Models
- Trainability-Oriented Hybrid Quantum Regression via Geometric Preconditioning and Curriculum Optimization
- Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models
- In-Hospital Stroke Risk-State Classification from PPG-Derived Hemodynamic Features
- Machine Learning-Based Classification of Jhana Advanced Concentrative Absorption Meditation Using 7 Tesla Functional Magnetic Resonance Imaging
- Decomposing Discrimination: Causal Mediation Analysis for AI-Driven Credit Decisions
- Are Independently Estimated View Uncertainties Comparable? Unified Routing for Trusted Multi-View Classification
- SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
- The General Theory of Localization Methods
- On the Residual Scaling of Looped Transformers: Stability and Transferability
- Representing and Detecting Label Ambiguity in IMU-Based Exercise Evaluation
- The C-index illusion: discrimination without calibration in published survival models
- DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation
- Diffract: Spectral View of LLM Domain Adaptation
- Learning Generalizable Reconstruction of High-Dimensional Neural Dynamics
- Multi-Source Wasserstein Distributionally Robust Graph Learning
- TDDM-Melatt: A Decoupled Memory and Diffusion Framework for Generalizable Encrypted Traffic Classification
- SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
- When Minute-Resolution Monitoring Meets Session-Level Injury Labels: Landmark-Based Discrimination in Elite Women's Football
- Distilled Continuous Diffusion Language Models Can Write Code in Few Steps---or One
- The BatchNorm Illusion: Diagnosing Normalization Artifacts in Machine Unlearning Evaluation
- Code-to-Harness: Distilling Black-Box Optimizers from Self-Play
- Byzantine-Robust Federated Fire Detection with a Rotating Coordinator
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization
- Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
- Synthetic Blips: Generalizing Synthetic Controls for Dynamic Treatment Effects
- Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution
- Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
- A Generalized Tangent Approximation based Variational Inference Framework for Strongly Super-Gaussian Likelihoods
- Can SGD Select Good Fishermen? Local Convergence under Self-Selection Biases
- Scalable Krylov Subspace Methods for Generalized Mixed-Effects Models with Crossed Random Effects
- On Universality of Non-Separable Approximate Message Passing Algorithms
- Limits of LLM Text Detectors in Education
- Nonlinear Dimensionality Reduction Techniques for Bayesian Optimization
- Bias-Corrected Data Synthesis for Imbalanced Learning
- Classical and quantum kernel fusion for two-sample testing
- Deep learning methods for inverse problems using connections between proximal operators and Hamilton-Jacobi equations
- PACEvolve: Enabling Progress-Aware Consistent Evolution
- AnyView: Synthesizing Any Novel View in Dynamic Scenes
- LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics
- Block-Norm Geometries for Online Mirror Descent with Sparse Losses
- An Empirical Markov Chain Car-Following (MC-CF) Model
- Exploring Urban Land Use Patterns by Pattern Mining and Unsupervised Learning
- Anomaly-Preference Image Generation
- Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization
- Independent Learning of Nash Equilibria in Partially Observable Markov Potential Games with Decoupled Dynamics
- Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model
- SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization
- GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science
- State-specific respiratory signatures for affective and stress recognition: Interpretable respiratory markers, autocorrelation lags, and compact CNN models
- Diffusion learning reveals viable parameter manifolds and compensation geometry in biological dynamical systems
- When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents
- Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
- From Information to Delegation: Mapping Human-AI Financial Decision Making
- AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
- Improving precipitation forecasts in an AI weather model using observational data
- EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments
- Decomposing LLM-Judge Uncertainty to Target Expert Labels
- LatentMD: Benchmarking Markdown Boundary Failures in LLM-Generated Text
- HuRo: Robotizing Human Videos for Scalable VLA Pretraining
- A Case-Bundle Operating Model for Coding Agents in OpenFOAM-Based CFD
- Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents
- Investigating Developer-Reported Software Security Testing Challenges
- Test-Driven Approaches to Software Engineering with Large Language Models: A Survey of Phases, Tasks, and Agent Skills
- Missing Dimensions: Integrating Human and Social Systems into Digital Twin Engineering
- Open Source Stewardship Communities: "We need you, but not your pull request"
- PQLS: A High-Performance Python Library for Steady-State Simulation of Open Quantum Systems
- Hieronym: Leveraging Hierarchical Multi-Source Information for Function Renaming in Stripped Binary
- A Retrieval-Augmented Automated Stakeholder for Requirements Elicitation Education: A Comparative Study
- Detecting HTTP Status Code Misuses in REST APIs via Static and Dynamic Analysis
- Intelligent Semantic Matching (ISM) for Video Tutorial Search using Transformer Models
- Beyond Establishing the Four-Day Workweek: Understanding Adaptation and Long-Term Survival in an Agile Software Organization
- MaRDMO: FAIR Documentation of In-Silico Research
- Scan the Skill, Govern the Action: Composing Registry Verdicts with Runtime Consequence Control
- Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis
- Mission Performance: Automatic and Adaptive Race Pace Progression for Autonomous Racing
- The case against JPEG XL
- Anecdotally, programmers dislike "reduce"
- What blog posts influenced your thinking the most?
- The contagion of fear
- Mergiraf: A syntax-aware git merge driver for a growing collection of programming languages and file formats
- Why Am I Still Programming
- How can you not be romantic about UNIX domain sockets?
- Why is the x86 undefined instruction called ud2? Why 2?
- Finished aerial maps in under 30 minutes
- Writing a Guix service from scratch, as a beginner
- We are all Product Engineers now
- There are only twelve 4x4 sudokus - and a cool trick for finding minimal subsets
- Watch what you say: Apple opens the door to a nightmare world of always-listening tech
- what if my git host were a static site generator?
- What are you doing this week?
- Switching to GNU Guix: A Beginner's Perspective
- I wish you the best in the Offline
- Purely Functional Operating Systems
- Homebrew 7.0.0
- A Letter from a Machine Learning Engineer
- Golang developers should try Odin
- Sticking Functions Where They Donʼt Belong
- Singeli: High-level interface for low-level programming
- This PCB is brought to you by Fable 5
- SplitFT: Fault Tolerance for Disaggregated Datacenters via Remote Memory Logging
- Why Data Quality Matters When Building Reliable AIoT Systems (いいね相当スコア: 0)
- The Power Platform Governance Repo - Standards Reviews and Inventory in Git (いいね相当スコア: 0)
- "Desafiando el futuro: استكشاف تقاطعات Web3, Automatización SaaS y Innovación en IA" (いいね相当スコア: 0)
- Fintech Innovation 2026: Essential Portfolio Access (いいね相当スコア: 0)
- Free 1M-context LLM API (Qwen3.8-Max, $0 Forever) (いいね相当スコア: 0)
- Best AC Repair in Lucknow, Mumbai, Delhi, Pune, Ahmedabad | Galaxy AC Service Center (いいね相当スコア: 0)
- I published a benchmark that was entirely my own bug (いいね相当スコア: 0)
- Creating a chatbot using a small local LLM (いいね相当スコア: 0)
- Fine Tuning a Local LLM to Categorize Questions (いいね相当スコア: 0)
- OpenAI Agents API Guide: How to Use It, Best Prompts & Use Cases (2026) (いいね相当スコア: 0)
- One API Key, Every Model: Calling GPT-6, Claude, Gemini, and DeepSeek Through a Single Endpoint (いいね相当スコア: 0)
- Comparing RAG to continued pretraining of LLMs (いいね相当スコア: 0)
- Teaching a local LLM a new domain (いいね相当スコア: 0)
- Changes to LLM pricing: Baidu, CoreWeave, Inceptron, Phala and Tencent (いいね相当スコア: 0)
- Changes to LLM pricing: Baidu, DigitalOcean, Inceptron and StreamLake (いいね相当スコア: 0)
- I Added More AI Agents to the Problem. Nothing Changed. (いいね相当スコア: 0)
- Your prompt has more writers than your context budget knows about (いいね相当スコア: 0)
- DeepSeek V4.1 Flash API Cost: Matches V4 Pro at 3-6x Less per Answer (いいね相当スコア: 0)
- How to Write an AI Agent Skill File That Reduces Guesswork (いいね相当スコア: 0)
- vLLM Preemption: Why 1 in 50 Requests Restarts From Scratch (いいね相当スコア: 0)
- I gave Claude Fable 5.1 and GPT 6 Astra control of my Robot Arm. Which do you think painted better?
- I asked Claude to build an operating system from scratch. A few days later it was running on a real laptop
- Simulation: what if you could throw anything into a black hole?
- Anthropic says Russian and Chinese threat actors used Claude for drone-swarm software and other weapons work
- Day 12ish of making a cozy game with no game dev experience
- I don't see how to upskill any further
- I am a CMU professor. This summer I taught a cohort of non-programmers to build and ship with Claude. 28 projects later, the fall cohort just started.
- AI doesn't just help me do, it also helps me think
- Anthropic gives you a month of limitless Fable 5.1, wwyd?
- Back to normal limit
- Claude Code May–August 2026 weekly limits promotion
- Sol Low managing Opus Medium - “I’m not publishing that”
- Claude Desktop and Fable 5.1 safeguards are still terrible, RT shader work in UnrealEngine flagged as [cybersecurity] work, and dropped to Opus 4.8
- My Max 20x ends today, as Claude code gets usagenerfed. With Code nerfed, will codex 20x usage be way ahead?
- I am SO OVER the its draining too quickly posts
- Built a cool way to visualize your Claude Code / Codex history
- Sharing my favorite Claude Code tips
- My Current Favorite Prompt: "I am frustrated with how you are doing."
- Built with Claude Code: a Discord channel where every thread is its own Claude Code session, so I can run it from my phone
- About the AI race ask to stop: I don’t want smarter models
- how do you keep long-running AI agents from quietly going off the rails?
- Built a bit-exact performance fork of a 25-year-old C++ emulator with Claude: +61% faster, 3 days, and a correctness gate proving the output never changed
- Are models much better than what they seem to with current agents?
- Weekly Self Promotion Thread
- We have achieved artificial middle management.
- OpenAI is building Codex Replay, a tool that invites Claude Code users to put Codex head-to-head on their own work — rerunning imported tasks and comparing the results
- GPT-5.6 Luna vs GPT-6 Astra: benchmark on 50 real PRs, looking for feedback on the methodology
- How do you get Astra to do less?
- A skill that does your app store research
- My review agent audited the wrong subsystem for $1.35 because of a 3-month-old session id
- Turned my SOL --> LUNA workflow into a repo-native Codex methodology
- Multi-agent coordination for prod dev
- Roblox Studio won’t connect to ChatGPT —Codex Computer Use detects zero Windows apps
- Git setup that keeps your AI agent from reading your whole codebase
- Need an app but no coding experience ?
- [Bi-Weekly Megathread] Project Showcase
- UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh
- For the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index.
- RTX PRO 5500 Blackwell (84GB) released
- Xi promotes open source AI zone among BRICS countries
- The new k2 horizon models seem like an absolute beast
- DeepSeek V4.1 Flash beats Astra on AA's new benchmark
- llama: add Maple 20B-A1B ternary MoE architecture (CPU) by AlexGabbia · Pull Request #27000 · ggml-org/llama.cpp
- Right to Intelligence. Protect your right to run local AI.
- 5090 Stock is Almost Gone
- K2 Horizon lineup is out on AA, and once again AA plots are misleading.
- Are there any organizations that are lobbying in favor of open source AI?
- 3k$ 128GB VRAM + 256GB RAM DDR4 Server
- Will we always have to rely on companies with the funds and resources to give us open models or can/will it be possible to democratize training for models capable of performing at or near the same level as the big closed ones in the future at some point?
- What actually makes you trust a local coding agent enough to leave it running unattended?
- Coding Agent running entirely in the browser with Pi + MiniCPM5-2B (webGPU)
- Is there a better small model than Qwen3.5 4B for a fast local AI assistant?
- Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors
- What are the current best retail GPUs for max VRAM at a reasonable price?
- Interesting Video by Asionometry on the state of SF Chips
- Intern-S2-397B (multimodal, reasoning, coding, and scientific agent capabilities)
- The Nvidia CMP 170HX -- 8GB -> 64GB ~ 1.49 TB/s
- Another Qwen3.8-27b Appreciation Post
- Micron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x
- Weekly Hiring Thread
- 15ms at P50 memory retrieval does absolutely nothing for a voice agent
- What’s one task you’ve actually stopped doing because of AI?
- The Memory Trust Gap: Why your agent acts on stale memory (and how to catch it)
- A good prompt is not a workflow. What turns an AI task into a repeatable process?
- Is anyone else finding agent harnesses less efficient than just using one strong agent + an orchestrator?
- How do you document AI generated code when nobody knows why it exists?
- I used AI tools to take our jewellery brand from 4 to 30 product videos a month. Sales grew nearly 10%
- Claude vs ChatGPT vs Gemini
- Complete beginner in AI automation: where should I start?
- Anybody solving authentication and authorisation for agents?
- Is git diff still the right thing for humans to review when coding agents work faster than we can?
- What building an agent harness around GPT-3.5 Turbo taught me: every fix was code, not a prompt
- What should a coding agent be allowed to do?
- Cold Calls experience
- I built a Open Source Research Map of Dead Science, Failed Companies & Cancelled Megaprojects
- The most expensive bug I hit building an agent wasn't in the code
- Title: The gap in AI agent builds isn't technical skill. It's that almost nobody here talks about who actually pays for this.
- What's a realistic monthly income if you actually know your field but hate selling yourself?
- Looking for a good general-purpose AI agent platform
- Testing frameworks for shared skills
- Using agents for real problems
- Parellel agents running tests on a single database
- AI Bot For Homeschool
- All they wanna do is get the girl
- China's Xi Jinping proposes BRICS 'open-source AI zone'
- GPT-6 Astra solves Portal, Baba Is You and other puzzles
- There's power in a name
- I made a news app that's just about progress and kindness in the world
- GPT 6 Astra playing Anno 117: Pax Romana from scratch
- File reading issue
- I have to ask Astra vs Sol
- What can we do to educate people on AI
- Are young engineers and recent graduates finding jobs?
- OpenAI now has an official Custom GPT retirement FAQ — Plugins are the recommended replacement
- Codex weekly usage dropping fast. What's wrong?
- I built a motion studio for coding agents. Here’s a 18-second Notion film made with it
- DeepSeek is ruthless
- The Doomsday Moat: Why AI’s Billionaires Want Washington to Stop the Clock
- Astra Significant Performance Degradation
- Got 6 Astra hallucinations
- "Some tools are temporarily unavailable" ... Is anyone else getting this on Sol High and their Work projects are now drunk as fuck?
- Any Idea when are we getting the $200 plan back for ChatGPT?
- Ai Question of the day
- GPT-6 Astra in a nutshell
- Quoting Laurie Voss
- commit-rewriter 0.1
- shot-scraper 1.12
- Humanity’s Last Invention — Richard Socher of Recursive
- Slashy Assistant
- Image to ASCII
- appdesigns
- Marqly 6.0
- AppZapper 3000
- Chinese AI labs are tightening the gap with the US by just being more efficient - Fortune
- DeepSeek Cut Its KV-Cache HBM Need 75% and SSD Need 87.5%. Micron and Sandisk Investors Should Pay Attention - Yahoo Finance
- Run a $10,000 AI Model at Home Using Kimi K3 and DeepSeek V4 - geeky-gadgets.com
- Why DeepSeek-V4.1-Flash Is Such an Exciting Open Model Release - KDnuggets
- AI Call | OpenAI>Anthropic in Q3 2026; New ARR, 5-Factor Regression Models - Thursday @ 2pm ET - Hedgeye
- How U.S. Export Controls Taught Chinese AI Companies To Innovate - Forbes
- DeepSeek-V4.1-Flash Packs 552B Parameters With Efficient MoE Inference - HackerNoon
- DeepSeek V4.1 Flash: What "Black Technologies" Make It So Competitive That It (Almost) Pushed the Pro Version Out of the Market - 36 Kr
- Alibaba Cloud Launches DeepSeek-V4.1-Flash AI Platform - GuruFocus
- DeepSeek's New AI Model Spooked Samsung and SK Hynix Investors - Startup Fortune
- The Sequence Radar - Issue 932: Last Week in AI: DeepSeek V4.1-Flash, AlphaGenome Atlas, Meta Muse, and OpenAI’s Proposed Math Breakthrough - TheSequence | Jesus Rodriguez
- DeepSeek Routes All V4-Pro API Traffic to V4.1-Flash at Flash Rates - Pandaily
- How trustworthy are the biggest names in AI? - Cybernews
- How DeepSeek V4.1 Achieves a 437x Memory Footprint Drop - geeky-gadgets.com
- 8:1 Kr Brief | Moonshot AI Files Police Report Over Malicious Rumors Targeting Founder & Employees, 25 Fields Medal Laureates Warn Against AI Encroachment on Mathematics, Zhipu AI Secures $5 Billion Financing - 36 Kr
- DeepS V4 Flash Vision EXP AI Model Processes Images For $0.00008 - geeky-gadgets.com
- Musk's xAI asks court to block Minnesota AI 'nudification' law during appeal - Reuters
- NVIDIA and SpaceXAI Link Grok Expansion With Orbital Computing - The Futurum Group
- Omneky Plugin Now Available on the Grok Bot Marketplace - PR Newswire
- Musk's xAI asks court to block Minnesota AI 'nudification' law during appeal - TradingView
- Grok Build Explained: What It Is and How to Try It - BASENOR - Tesla Accessories
- Groq vs Grok 4.6: Spec, Price & Speed Compared [2026] - tech-insider.org
- The SpaceX-xAI Merger: 5 Numbers That Tell the Real Story - BASENOR - Tesla Accessories
- SpaceX AI Team to Build a Company Live with Grok Bot - varindia.com
- Grok picks 3 SpaceX ETFs to buy now - finbold.com
- ToolSearchは「無効化されたツール」を救わない: Claude CodeのTodoWrite境界 (いいね相当スコア: 0)
- Rust製AIエージェントフレームワーク入門、rigとAutoAgentsの選び方 (いいね相当スコア: 0)
- AIパートナーに「自律性」を持たせる:定期的にAIパートナーが自分の意思でお喋りしてくれるコードを作ってみた (いいね相当スコア: 0)
- ローカルLLMにレビューさせていた話の続報(前回の構成、実は動いていませんでした) (いいね相当スコア: 2)
- Instruction Tuningを多タスクSFTとして設計する (いいね相当スコア: 0)
- 10層で見るAIの現在地。エンジニアが磨く知識はどこにあるか (いいね相当スコア: 6)
- 【第1回】AIエージェントに実行させない——身体と頭脳の分離 (いいね相当スコア: 1)
- LLM費用¥0で公開リポジトリの保守を回してみた (いいね相当スコア: 2)
- 複数LLMを混ぜるのをやめた。理由は対照実験の数字 (いいね相当スコア: 0)
- ドキュメントをLLMで自動同期させたら、担当者交代のコストが変わった話 (いいね相当スコア: 0)
- 実際の企業コードでAIモデルの性能を測る「Real-SWE」とは? (いいね相当スコア: 0)
- LangChainエージェント入門 — LLMに「道具」を持たせて自律的に動かす (いいね相当スコア: 0)
- Skill を「プロンプト」と呼ぶのをやめた — 自然言語で書いたコードとして扱う (いいね相当スコア: 1)
- 57. MCPは誰が許可する? (いいね相当スコア: 0)
- Gemini 5モデル1509試行、注入を防ぐのは手数だった (いいね相当スコア: 2)
- Windows で Docker Desktop を使っているあなたへ:「WSL3」「Docker Desktop 不要」の噂を確かめる (いいね相当スコア: 0)
- Multi-Agentにおける「多様性」について考えてみた (いいね相当スコア: 1)
- AIエージェントが「該当なし」と言ったとき、確かめるべき3つのこと (いいね相当スコア: 4)
- 常駐型AIチャットエージェントのLLMプロバイダ障害対応・フェイルオーバー設計 (いいね相当スコア: 0)
- 依頼から成果物の保存までつなぐMacアプリ「Genie」を作った。LLMを組み込んで直した3つのこと (いいね相当スコア: 1)
- 事件簿:「頭にきた」が喜びになる ― 一語が全体を狂わせる ― KotobaCore開発秘話(4/8) (いいね相当スコア: 1)
- サーバー側処理なし! Blazor WebAssembly スタンドアロンアプリでベクトル検索を実装する (いいね相当スコア: 3)
- 競馬AIの検証コード46,000行と、合格を疑うコード0行 (いいね相当スコア: 0)
- その予測、当たる前提で使っていませんか——Conformal Predictionと、埋め込み距離という2つの見積もり方 (いいね相当スコア: 0)
- Co-purchaseからPersonalized Recommendationまで — 商品レコメンドの仕組みをシンプルに理解する (いいね相当スコア: 1)
- 個人実証:機微データに触れないログ異常検知は成立するか(型で強制する設計と実測精度) (いいね相当スコア: 1)
- 読書メモ『強化学習を学びたい人が最初に読む本』 (いいね相当スコア: 13)
- ESP を自分で測ってみる (いいね相当スコア: 2)
- 重みは学習するのか、測るのか。ハエの脳が問い直すAIのつくり方 (いいね相当スコア: 89)
- 暴力的研究開発:安くなった探索と、マルチエージェントの判断力 (いいね相当スコア: 0)
- DeepSeek V4.1 Flash API コスト、V4 Pro 同等で 3-6x 安い (いいね相当スコア: 0)
- AIチャットボット公開前のセキュリティテスト設計:30項目を7層に分ける (いいね相当スコア: 0)
- 【書評】AWSではじめる生成AI ―RAGアプリケーション開発から、基盤モデルの微調整、マルチモーダルAI活用までを試して学ぶ (いいね相当スコア: 0)
- sklearn.datasetsで使えるデータセット・取得関数・生成関数一覧 (いいね相当スコア: 0)
- なぜSkynetは生まれたのか — Bengioの「AIエージェントはなぜ嘘をつき、裏切り、結託するのか」から考える (いいね相当スコア: 0)
- 【技術解説】隠れマルコフモデルで金融市場のレジームチェンジを予測するPython実装ガイド (いいね相当スコア: 0)
- トランスフォーマーは早すぎた (いいね相当スコア: 0)
- 🌇 Day58|AIはなぜ同じ質問でも違う回答をする?「確率」で2分理解 (いいね相当スコア: 取得失敗)
- AIは賢いのに、なぜ締切を忘れるのか 一人会社に「週次レビュー」を実装した話 (いいね相当スコア: 取得失敗)
- AIエージェント50体で会社を回す私が、それでも売上を取りこぼしていた話 (いいね相当スコア: 取得失敗)
- 生成しないのに AI モデル? エンコーダ型・埋め込みモデルの歴史と用途を word2vec から 2026 年まで一本でたどる (いいね相当スコア: 取得失敗)
- 【第4章:LLMの正体編】第30話:「Attentionだけじゃなかったの!?」Transformer Blockを支えるFeed Forward・Residual・Normalization (いいね相当スコア: 取得失敗)
- 【AIってネットなしでも動くの?】「ローカルLLM」を初心者向けに整理 (いいね相当スコア: 取得失敗)
- AIオーケストレーションの資料を読んで学んだことの備忘録 (いいね相当スコア: 取得失敗)
- AI研究は止まらない (いいね相当スコア: 取得失敗)
- 同じ「GPT-5.6」だと思って課金を切った話 ——カスタム指示は、載せた時点では効いていない (いいね相当スコア: 取得失敗)
- 【月200時間の泥臭い業務が0秒に】メール仕分け・問い合わせ対応・レポート作成を完全自律化する「Superpowers × MCP × LangGraph」次世代AIエージェント導入完全ロードマップ (いいね相当スコア: 取得失敗)
- 【年間1,200時間の泥臭い残業が0秒に】設計・購買・製造の表記ズレや旧型番を100%完全自動で紐付ける「Polars × Dynamic Graph × Instructor」全自動BOM突合パイプライン構築全書 (いいね相当スコア: 取得失敗)
- Qwen3.8-Flash-Next:UD-IQ3_XXSにp5.jsで「夕暮れのピクセルローカル線」を書かせてみた (いいね相当スコア: 取得失敗)
- AIに何度言っても指示を守らない…原因と対策を解説 (いいね相当スコア: 取得失敗)
- 【生成AIニュース+】『qwen38-flash-next-recipe』『StepAudio 3 Gen』『Comfyui-Auk-T8』『Comfyui-Image-Stitch』『SeeSee Civitai Metadata』『MiniMax H3 TaoMate 3step LoRA』『BUNNY H3 Conditioning Bridge』『ComfyUI_toyxyz_test_nodes』『ComfyUI-Multiline-Text-Prompt』『AnimalLift』他 (いいね相当スコア: 取得失敗)
- AIエージェントに必要なのは、賢い頭より「壊していい作業場」かもしれない (いいね相当スコア: 取得失敗)
- 【Termux活用】スマホの限界に挑む!手のひらサイズでAIを動かす試み (いいね相当スコア: 取得失敗)
- 今更聞けないAI用語|RAGとファインチューニングは何が違う? (いいね相当スコア: 取得失敗)
- 現在のAI制御の壁は、人間との対話の難しさと「同じ」であることの証明(*特許出願済) (いいね相当スコア: 取得失敗)
- 毎朝AIラボ #241|2026-09-14 (いいね相当スコア: 取得失敗)
- 【保存版】AIはどう作られ、どう答え、どう仕事をするのか?——20枚の図でたどる、用語と仕組みの入門(全19章) (いいね相当スコア: 取得失敗)
- GeminiSparkで試す最新論文解読:AIパートナーと『NCP(Next Concept Prediction)』を深掘りしてみた (いいね相当スコア: 取得失敗)
- Next.js受託でLLMに会話履歴をどこまで渡すか決める5つの判断 (いいね相当スコア: 取得失敗)
- Fusionの最小半径解析、工具選びの見落としを色で防ぐ|Fusion CAD (いいね相当スコア: 取得失敗)
- Fusion CAMの複数WCS、同じ加工をまとめる時短術|Fusion CAM (いいね相当スコア: 取得失敗)
- AnthropicのMCP共同作者が来日、基調講演で語るこれからのMCP注力領域。AGNTCon+MCPCon Japan 2026
- OpenAIやAnthropicなどAIベンダごとのAPIの違いを吸収し統合する「Agent Router」、Linux Foundation傘下で業界標準へ
- マイクロソフトがRust言語をC++/C#/TSに並ぶ社内のTier 1言語にしたことを明らかに。Windowsネイティブな社内の開発環境と統合
- AIのせいでエンジニアの75%を解雇したCSSフレームワークのTailwind、Shopifyによる買収を発表。今後も安定的な開発を維持すると
- TailscaleのVPNにAIエージェントを組み込める「Aperture」正式リリース。AIによるTailscaleやノードの操作も可能に
- AIの成果物を人が受け取るための設計ワークフロー
- AIとGrafana Foundation SDKで進めるGrafanaダッシュボードのIaC化
- Local LLMを社内に提供! Local LLM Model as a Serviceとその取り組みについて