James Trappett

Paper Reviews, Security News & Technical Analysis

Latest Articles

LLM Agents Run Controlled Experiments via Simulation Models

27 August 2026

A multi-agent LLM framework conducts controlled simulation experiments for pharmaceutical process design, outperforming language-only reasoning on specificity a

LLM AgentsMulti-Agent SystemsScientific AISimulationProcess Optimization

RENDER: How Evidence Formatting Skews LLM Memory Benchmarks

27 August 2026

RENDER shows that how memory evidence is formatted for LLMs changes accuracy by up to 73 points, exposing a critical uncontrolled variable in RAG evaluation.

LLM EvaluationRAGMemory SystemsBenchmarkingNLP

Weibull Weight-Scale Growth Predicted by Corpus Entropy

27 August 2026

A pre-training data statistic, bigram conditional entropy, predicts transformer weight-scale growth via a power-law relation with R²=0.941 across 23 training ru

TransformersTraining DynamicsInformation TheoryWeight AnalysisML Theory

Agentic Scaffolding Amplifies Sycophancy in LLMs

26 August 2026

New research shows agentic feedback loops systematically increase sycophantic behaviour in LLMs, causing a mean 6.3pp accuracy drop across six frontier models.

LLM SafetySycophancyAgentic AIAI AlignmentNLP Research

KVBoost: Chunk-Level KV Cache Reuse for Faster LLM Inference

26 August 2026

KVBoost achieves 4.49x faster LLM prefill by reusing KV cache chunks at arbitrary prompt positions, outperforming vLLM prefix caching by 16% with no quality los

LLM InferenceKV CacheTransformer OptimizationSystems ML

Model Collapse in Generative AI: Causes and Countermeasures

26 August 2026

A systematic review of model collapse in generative AI, covering causes, observable signals, and mitigation strategies for self-consuming training loops.

Generative AIModel CollapseSynthetic DataLLMsDiffusion Models

BF1: Sparse Attention Retrofit for Long-Context Transformers

25 August 2026

BF1 achieves 10.91x prefill speedup at 32K tokens via dyadic sparse attention, with better perplexity than dense continued training after selective retrofit.

Efficient TransformersSparse AttentionLong ContextInference OptimizationNLP

Clinical Lost-in-the-Middle: Positional Bias in EHR LLMs

25 August 2026

New research characterises positional retrieval bias in clinical EHR processing and introduces QCCS, a lightweight query-conditioned context selection method th

Clinical NLPLLMsRetrieval-Augmented GenerationEHRLong-Context

LLM Safety Gaps: Detecting Harmful Intent in Early Layers

25 August 2026

New research shows LLM safety alignment fails against semantic camouflage attacks, but early-layer probing detects harmful intent with 20-50% better accuracy.

AI SafetyLLMAdversarial MLMechanistic InterpretabilityResearch Review

Russian Backdoor in Slovak Speed Cameras: A Supply Chain Case Study

24 August 2026

Slovakia's NBU found SMS-triggered backdoors in Russian-made traffic cameras. A technical breakdown of the vulnerabilities and supply chain security implication

CybersecurityIoT SecuritySupply ChainCritical InfrastructureEastern Europe

English to Claudish: Analysing LLM Refusal Pattern Translators

23 August 2026

A critical analysis of the English-to-Claudish translator tool, examining what it reveals about Claude's refusal behaviours, alignment mechanisms, and the broad

AI SafetyLLM ResearchAnthropicPrompt EngineeringAI Alignment

NanoGPT Speedrun: Benchmarking Frontier AI Agents at Scale

23 August 2026

Prime Intellect ran 153 autonomous agent runs across 18 frontier models on the nanoGPT optimizer speedrun. Here is a technical breakdown of what the results rev

AI AgentsBenchmarkingLLM ResearchAutonomous SystemsMachine Learning
View all 278 articles →