Template-Type: ReDIF-Paper 1.0 Title: Pearson residuals in logistic regression: comparison between logit and glm Author-Name: Marta Ponzano Author-Workplace-Name: Department of Life Sciences, Health and Health Professions, Link Campus University Author-Name: Rino Bellocco Author-Workplace-Name: Department of Statistics and Quantitative Methods, University of Milano?Bicocca, Milan, Italy; 3Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Solna, Sweden. Abstract: Logistic regression can be used to model the relationship between a set of covariates and a binary outcome. The predicted probabilities from the estimated model are inherently identical for observations sharing the same covariate pattern. In Stata, logistic regression can be performed using the logit command and the glm command. Both procedures calculate Pearson residuals, but the underlying computation differs substantially. In logit, Pearson residuals are calculated at the level of covariate patterns, and the Pearson chi-square goodness-of-fit test statistic is obtained by summing their squared values across covariate patterns. In contrast, in glm, Pearson residuals are computed at the individual level, and the Pearson deviance is defined as the sum of their squares over all observations. Unlike logit residuals, glm residuals can potentially vary among observations sharing the same covariate pattern, depending on the observed outcomes. Unlike logit residuals, glm residuals can potentially vary among observations sharing the same covariate pattern, depending on the observed outcomes. In general, except in the case of unique covariate patterns, Pearson residuals from logit and glm are not therefore the same. To illustrate these key differences, we present an example under three distinct scenarios: a single continuous covariate, a single categorical covariate, and a set of covariates of different types. Given the importance of post-modeling estimation, it is crucial to be aware of how Pearson residuals are calculated and to recognize that the Pearson goodness-of-fit test relies on residuals computed at the covariate pattern level. We hope this contribution supports the correct interpretation and practical use of Pearson residuals in Stata. Creation-Date: 20261001 Handle: RePEc:boc:neur26:06 Template-Type: ReDIF-Paper 1.0 Title: Autonomous research agents for mathematical conjecture testing: Bridging Stata 19 and Agentic AI Author-Name: Prasad Kothari Abstract: As Large Language Models (LLMs) transition from generative chatbots to autonomous reasoning agents, the potential for automating complex scientific discovery workflows has expanded. This presentation introduces an agentic AI framework designed to solve and verify mathematical conjectures - such as those in extremal combinatorics and geometric structures - by leveraging Stata 19 as a primary engine for rigorous statistical validation and structural deep learning. This session demonstrates how an autonomous agent can: orchestrate workflows using the ReAct paradigm to decompose high-level mathematical hypotheses into executable Stata code; perform statistical verification by employing Stata's Bayesian variable selection (bayesselect) and Structural Equation Modeling (SEM) to test the stability of generated conjectures against large-scale synthetic datasets; and carry out topological data analysis (TDA) by integrating external Python-based TDA libraries with Stata's visualization tools to identify geometric patterns in mathematical objects. The session will include a "Tips and Tricks" segment on building "Statistics Agent Skills" - local markdown-based instruction sets that allow LLMs to maintain context-aware modeling strategies within the Stata environment. We demonstrate that by embedding Stata's rigorous econometric standards into an agent's "second brain," we can mitigate the logical reasoning failures common in generic LLMs while accelerating the pace of decentralized science. Creation-Date: 20261001 Handle: RePEc:boc:neur26:07 Template-Type: ReDIF-Paper 1.0 Title: Approximate kernel regression in Stata using random Fourier features Author-Name: Christopher Rose Author-Workplace-Name: Norwegian Institute of Public Health Abstract: Nonlinear regression has wide application, including adjusting for confounders and secular trends in observational studies, modeling dose- and exposure-response relationships, predicting risk, and characterizing treatment effect heterogeneity. Kernel methods facilitate flexible nonparametric regression, but the computation time and memory of classical approaches scale quadratically in the number of observations. In this talk, I will describe the makerff command for approximating kernel regression using trigonometric basis functions called random Fourier features. With this approximation, time and memory scale linearly in the number of observations for a given number of features. The command generates these features from a covariate varlist; kernel regression can then be approximated by ridge-penalized regression for continuous, binary, count, and time-to-event outcomes. I will present examples in which the command is used to approximate kernel logistic regression for risk prediction, and kernel Cox regression with g-computation and bootstrap inference for marginal effect estimation. Creation-Date: 20261001 Handle: RePEc:boc:neur26:08 Template-Type: ReDIF-Paper 1.0 Title: Introduction to machine learning and AI using Stata Author-Name: Chuck Huber Author-Workplace-Name: StataCorp Abstract: This talk will briefly review the concepts and jargon of machine learning (ML) and artificial intelligence (AI), and demonstrate how to use these tools in Stata. Specific examples will include random forests and gradient boosting machines using Stata's suite of h2oml commands, and the user-written commands "chatgpt", "claude", "gemini", and "grok". Creation-Date: 20261001 Handle: RePEc:boc:neur26:09 Template-Type: ReDIF-Paper 1.0 Title: Projecting cancer prevalence through combining age-period-cohort models and survival models Author-Name: Paul Lambert Author-Workplace-Name: Cancer Registry of Norway, Norwegian Institute of Public Health Abstract: The number of people living with a diagnosis of cancer has increased over the last decades due to increasing incidence, improved survival, population growth and a shifting age distribution. Projections of cancer incidence and prevalence are used to estimate the future burden of cancer and help inform future resource requirements. Prevalence projections can be estimated by combining statistical models to predict future incidence and future survival, alongside estimates of future population structure. I will describe an approach, together with new commands, to predict future cancer prevalence that combines age-period-cohort (APC) models incorporating natural splines and flexible parametric survival models. For the APC models different assumptions about future incidence rates can be made through including various combinations of different link functions, moving the upper boundary knot of the spline function and 'dampening' future incidence. Different approaches to extrapolating future survival can also be made through use of period analysis and/or careful modeling of calendar time. As with any extrapolations a well-informed sensitivity analysis is vital, as well as investigation of realistic "what if?" scenarios. I will describe new Stata commands to fit the APC models (apcmodel) and a set of postestimation commands to predict future incidence to combine the APC model with a survival model fitted using stpm3. Creation-Date: 20261001 Handle: RePEc:boc:neur26:10 Template-Type: ReDIF-Paper 1.0 Title: Reference adjusted survival measures as an alternative to net survival in population-based cancer studies Author-Name: Rebecka Johansson Author-Workplace-Name: Department of Medical Epidemiology and Biostatistics, Karolinska Institutet Author-Name: Therese M-L Andersson Author-Workplace-Name: Department of Medical Epidemiology and Biostatistics, Karolinska Institutet Author-Name: Paul Lambert Author-Workplace-Name: Cancer Registry of Norway, Norwegian Institute of Public Health Abstract: Age-standardized marginal net survival is the standard measure for comparing cancer survival across populations, but it is difficult to communicate to patients and clinicians. Reference adjustment has been proposed as an alternative, providing all-cause survival and crude probabilities of death that are both interpretable and comparable across populations by standardizing other-cause mortality to a common reference. In Stata non-parametric net and reference adjusted survival are available using the stpp command. A parametric flexible parametric model-based approach is possible using the stpm3 command followed by the postestimation standsurv command. I will demonstrate the different approaches using data from the Swedish Cancer Register on colon cancer, lung cancer and melanoma. In addition I will describe a simulation study following the ADEMP framework, conducted to assess the statistical properties of net survival and reference adjustment under controlled conditions, including their bias, coverage and power to detect differences between populations, as well as their robustness to errors in the population mortality file. The simulation study demonstrates two main findings. First, non-parametric reference adjustment shows substantially higher power than net survival to detect differences between populations for long-term survival estimates, with the advantage increasing with follow-up time and being most pronounced for cancers with good prognosis. Second, reference adjustment is substantially more robust to proportional errors in the population mortality file than net survival, a consequence of the mathematical structure of the estimator in which individual and reference mortality rates fully or partially cancel. Together these findings suggest that reference adjustment could be a valuable alternative to net survival in population-based cancer research. Creation-Date: 20261001 Handle: RePEc:boc:neur26:11 Template-Type: ReDIF-Paper 1.0 Title: Causal mediation analysis with multiple mediators Author-Name: Joerg Luedicke Author-Workplace-Name: StataCorp Abstract: Causal mediation analysis allows for the decomposition into direct and indirect causal effects. This form of decomposition provides a fine-grained look at the pathways through which causal effects operate. In this presentation, we will review the basic principles of causal mediation in the context of the potential-outcomes framework. In addition, we will discuss some of the intricacies that we face when having more than one mediator variable, and we will discuss a number of examples with two mediators to illustrate the corresponding challenges. Creation-Date: 20261001 Handle: RePEc:boc:neur26:12 Template-Type: ReDIF-Paper 1.0 Title: Running multiple instances of Stata using the sessions command Author-Name: Paul Lambert Author-Workplace-Name: Cancer Registry of Norway, Norwegian Institute of Public Health Abstract: I will describe the sessions command, which is a simple tool to run "embarrassingly parallel" Stata sessions. The sessions command creates multiple Stata sessions with the ability to pass the same or different Do files to each session, with separate arguments for each session. Before starting a new Stata session, the sessions command will check whether there is sufficient free memory, the CPU load is not too high, and that the maximum number of instances of Stata (set by the user) is not exceeded. A list of the Stata sessions that are queued, currently running, and have completed is updated every few seconds. The sessions command allows the user to wait until specific Do files have completed before continuing, for example all analysis files may need to be completed before producing tables/graphs of results. Creation-Date: 20261001 Handle: RePEc:boc:neur26:13 Template-Type: ReDIF-Paper 1.0 Title: Bias-correction in adaptive trials using Stata Author-Name: Christopher Rose Author-Workplace-Name: Norwegian Institute of Public Health Abstract: Adaptive trials permit design features - sample size, randomization ratios, or enrolment criteria - to be changed based on accumulating data and are of particular interest in drug trials. A novel potential application is in public health and social measures (PHSM) trials. While PHSM were widely used during COVID-19, evidence on their benefits and harms remains limited. PHSM trials are challenging because the sampling frame is confined to the winter respiratory infection season, and measures can be burdensome, limiting enrolment. Adaptive PHSM trials can reduce sample sizes, spread recruitment across multiple winters, and allow early stopping for efficacy or futility, facilitating reallocation of research resources. While adaptive designs control type I and type II error, data-dependent adaptations tend to bias conventional estimators. This is well-documented, and a substantial corrective literature exists, yet systematic reviews find that bias correction is rarely used. This may be because analytic corrections are design- and outcome-specific, and there is little software support. I will present a prototype Stata command for bias correction applicable to arbitrarily complex adaptive designs. I will also show results for simulation-based experiments validating the command in a range of adaptive designs, including a real PHSM trial: an ongoing group sequential trial of portable air purifiers in schools to reduce student absence due to illness. Creation-Date: 20261001 Handle: RePEc:boc:neur26:14 Template-Type: ReDIF-Paper 1.0 Title: datamirror: coefficient-preserving synthetic data for restricted microdata Author-Name: Jeffrey Clark Author-Workplace-Name: Department of Economics, Stockholm University Abstract: Reproducibility has become a condition of publication, but studies on restricted microdata remain its standing exception. The code can leave the secure environment; the data cannot. Synthetic data is the natural substitute, yet existing generators preserve distributions, not the published estimates. Re-run a published regression on their output and a different number comes back. This talk introduces datamirror, a Stata command that adds a coefficient-preserving layer to distributional synthesis. Inside the enclave the researcher checkpoints a chosen set of regressions, and the extract step writes a small bundle of aggregates with no individual records. From that bundle alone, on any machine, anyone can rebuild the synthetic data, and the checkpointed regressions return their published coefficients, exactly for linear estimators and within sampling noise for the rest. No synthetic-data tool does this. The same bundle is both a replication package and a synthetic copy to work on outside the enclave. Across four American Economic Association replication packages, datamirror reproduces 377 of the 378 checkpointed regressions that yield an evaluable coefficient. The talk includes a live demonstration of the checkpoint, extract, and rebuild workflow. Creation-Date: 20261001 Handle: RePEc:boc:neur26:15