Indiana Prime Time Third Grade Achievement Evaluation Data
Source:R/data_prime_time_achievement.R
prime_time_achievement.RdThe complete student level data file from the 2000 to 2001
Indiana Department of Education program evaluation of Project
Prime Time, reported in Lapsley, Daytner, Kelley, and Maxwell
(2002, ERIC ED466679). The evaluation examined the academic
performance of N = 10,927 third grade students in
n = 573 classrooms (here 586 by the
paste(corp, school, class) rule), n = 163
schools, n = 61 school corporations, and 9 Indiana
educational service regions as a function of class size, pupil
to teacher ratio, and the presence of a Prime Time
instructional assistant. The data have been used as a
multilevel example through chapters 3, 4, 6, 9, and 10 of
Finch, Bolin, and Kelley (2019, Multilevel Modeling Using
R, 2nd ed., CRC Press) and are made available here as a
benchmark data set for the design, measurement, and analysis of
nested data.
Format
A data frame with 10,927 observations on 113 variables.
Variables fall into seven blocks: three derived unique cluster
identifiers, student level demographics and ability and
achievement scores, classroom level variables, school level
variables, and school corporation (district) level variables.
Original Indiana DOE variable spellings are preserved, including
the typos calender (calendar), hispanc1 and
hispanc2 (Hispanic), and rmediate (remediate), for
code compatibility with Finch, Bolin, and Kelley (2019). The
SPSS variable label from the source file is available as
attr(prime_time_achievement$VAR, "label") for every
variable carried over from the SPSS file; the one derived recode
without a source label is classize (see its entry below).
idStudent identifier in the source file (not guaranteed unique; see
class_idetc. for stable cluster keys).regionIndiana educational service region, coded 1 to 9. Sampling stratifier (25% of corporations per region).
corpSchool corporation (district) numeric identifier. Note: corp 2400 appears in both region 2 (12 schools) and region 3 (1 school) in the source file, so
paste(region, corp)is the cleaner cluster key; seecorp_id.schoolNumeric school identifier (unique within corporation).
classClassroom number within school, 1 to 8. Not unique across schools.
corp_idDerived.
paste(region, corp, sep = "_"). 61 distinct values, matching the count of school corporations reported in Lapsley et al. (2002).school_idDerived.
paste(corp, school, sep = "_"). 163 distinct values.class_idDerived.
paste(corp, school, class, sep = "_"). 586 distinct values. (The published report counted 573 classrooms; the small discrepancy reflects a different counting convention used in the manuscript.)gender1 = Female, 2 = Male. 45
NA.ageStudent age in months.
raceIndiana DOE 6-category ethno-racial code: 1 = American Indian / Alaskan, 2 = African American, 3 = Asian American, 4 = Hispanic American, 5 = Caucasian American, 6 = Multi-racial.
gereadGates-MacGinitie reading.
gevocabGates-MacGinitie vocabulary.
gereadcmGates-MacGinitie reading composite.
gelangGates-MacGinitie language.
gelangmcGates-MacGinitie language mechanics.
gelangcmGates-MacGinitie language composite.
gemathGates-MacGinitie mathematics.
gemathcpGates-MacGinitie mathematics computation.
gemathcmGates-MacGinitie mathematics composite.
getotalGates-MacGinitie total.
ncreadNCE reading (ISTEP+).
ncvocabNCE vocabulary (ISTEP+).
ncreadcmNCE reading composite (ISTEP+).
nclangNCE language (ISTEP+).
nclangmcNCE language mechanics (ISTEP+).
nclangcmNCE language composite (ISTEP+).
ncmathNCE mathematics (ISTEP+).
ncmathcpNCE mathematics computation (ISTEP+).
ncmathcmNCE mathematics composite (ISTEP+).
nctotalNCE total composite (ISTEP+), the criterion variable in the Lapsley et al. (2002) HLM analyses.
aareadAANCE reading.
aavocabAANCE vocabulary.
aareadcmAANCE reading composite.
aalangAANCE language.
aalangmcAANCE language mechanics.
aalangcmAANCE language composite.
aamathAANCE mathematics.
aamathcpAANCE mathematics computation.
aamathcmAANCE mathematics composite.
aatotalAANCE total.
npanverbNPA nonverbal reasoning.
npamemNPA working memory.
npaverbNPA verbal reasoning.
npatotalNPA total.
csiCognitive Skills Index (student level).
multiMulti-age classroom indicator (1 = Yes, 2 = No).
typmultiType of multi-age classroom (1 = 1st-2nd-3rd grades, 2 = 2nd-3rd, 3 = 3rd-4th, 4 = 2nd-3rd- 4th;
NAwhenmulti == 2).clenrollOfficial class enrollment.
classizeProject STAR style class size category: 1 = small (roughly 12-17), 2 = regular (roughly 18-22), 3 = regular-larger (roughly 23-26), 4 = large (27 or more). Boundaries follow the STAR classification (Pate-Bain and Achilles, 1986).
ptratioClassroom pupil to teacher ratio (IDOE formula: enrollment / [1.00 per full time teacher + 0.33 per full time aide + 0.165 per part time aide]).
ptiaPrime Time Instructional Aide status: 1 = aide present, 2 = no aide, 3 = other assistant listed. The focal treatment indicator.
ptstatusStatus of Prime Time aide: 1 = full time in classroom, 2 = part time in classroom;
NAwhenptia != 1.localeNCES locale code (1 = large central city, 2 = mid-size central city, 3 = urban fringe of large city, 4 = urban fringe of mid-size city, 5 = large town, 6 = small town, 7 = rural).
chapter1School receives Title I (legacy "Chapter 1") money? 1 = Yes, 2 = No.
sesSchool SES, the IDOE percentage of students not eligible for subsidized lunch (0 to 100; higher = more affluent). The school level SES variable used in the Lapsley et al. (2002) HLM analyses.
contextIDOE contextual rank for the school.
calenderSchool calendar type (1 = traditional, 2 = year round). Typo preserved from source.
senrollBuilding (school) enrollment.
sattendBuilding attendance rate (percent).
white1School percent White.
black1School percent Black.
hispanc1School percent Hispanic (typo preserved).
asian1School percent Asian.
aindian1School percent American Indian.
multi1School percent multi-racial.
total1School total percent non-white.
noteachNumber of teachers in the building (full time equivalent).
avgage1School average teacher age.
avgexp1School average teacher experience (years).
avgsal1School average teacher salary (dollars).
spertSchool students per teacher.
thrdclssNumber of third grade classrooms in the building.
thrdstudNumber of third graders who took ISTEP+ in the building.
passla1Building percent passing language arts.
passmth1Building percent passing math.
passbth1Building percent passing both.
tmnnce1Building total battery mean NCE.
rmdnce1Building reading median NCE. Note: the 56 rows from one building (
corp5740,school6187) carry the source value 7603, an evident data-entry error in the Indiana DOE file (an NCE is on the 1 to 99 scale, and the same building's other median-NCE columns are in range). The value is preserved as shipped rather than silently corrected, since the true value cannot be recovered; drop or set it toNAbefore analyzing this column.lamdnce1Building language arts median NCE.
mmdnce1Building mathematics median NCE.
tmdnce1Building total battery median NCE.
avgcsi1Building average Cognitive Skills Index.
geogGeographic category of the corporation (1 = urban, 2 = suburban, 3 = town, 4 = rural). Sampling stratifier within region.
toteppCorporation total expense per pupil (1997–1999, dollars).
cenrollCorporation enrollment (all grades).
cattendCorporation attendance rate (percent).
freelnchCorporation percent eligible for free lunch.
lepCorporation percent with limited English proficiency.
specedCorporation percent in special education.
minorityCorporation percent minority.
white2Corporation total White public enrollment (raw count).
black2Corporation total Black public enrollment.
hispanc2Corporation total Hispanic public enrollment (typo preserved).
asian2Corporation total Asian public enrollment.
aindian2Corporation total American Indian public enrollment.
multi2Corporation total multi-racial public enrollment.
total2Corporation total non-white public enrollment.
thrdadmCorporation third grade ADM (average daily membership).
thrdtechCorporation third grade teachers.
avgage2Corporation average teacher age.
avgexp2Corporation average teacher experience (years).
avgsal2Corporation average teacher salary (dollars).
thrdaideCorporation third grade aides.
passla2Corporation percent passing language arts.
passmth2Corporation percent passing math.
passbth2Corporation percent passing both.
tmnnce2Corporation total battery mean NCE.
rmdnce2Corporation reading median NCE.
lamdnce2Corporation language arts median NCE.
mmdnce2Corporation mathematics median NCE.
tmdnce2Corporation total battery median NCE.
rmediateCorporation remediation funding per pupil (dollars). Typo preserved.
Source
Indiana Department of Education program evaluation of Project Prime Time, 2000 to 2001 academic year. Sample of 10,927 third grade students in 586 classrooms in 163 schools in 61 corporations in 9 educational service regions. The records are public data that the author, a member of the evaluation team, is authorized to distribute.
Lapsley, D. K., Daytner, K. M., Kelley, K., and Maxwell, S. E. (2002). Teacher aides, class size and academic achievement: A preliminary evaluation of Indiana's Prime Time. Paper presented at the Annual Meeting of the American Educational Research Association, New Orleans, LA, April 1-5, 2002. ERIC document ED466679.
Details
Study background. Indiana's Prime Time program, phased in beginning 1984 to 1985 (Indiana statute; House Bill 1166 of 2001 codified the modern funding formula), was one of the earliest state level initiatives in the United States to use a funding formula to reduce class size and pupil to teacher ratio in Kindergarten through third grade. Funds were distributed to school corporations to maintain a corporation average pupil to teacher ratio of 18:1 in K and grade 1 and 20:1 in grades 2 and 3; corporations could meet the target by hiring additional teachers or, more commonly, paraprofessional instructional assistants. Along with Tennessee's Project STAR (Pate-Bain and Achilles, 1986; covered by Education Week, Research: Sizing Up Small Classes, February 2001), Prime Time was a widely cited national model.
In 1999 the Indiana Department of Education funded a three year program evaluation of Prime Time. The third year of the evaluation, conducted by Daniel K. Lapsley and Katrina M. Daytner (Ball State University and Western Illinois University) with technical assistance from Ken Kelley and Scott E. Maxwell (University of Notre Dame), examined the academic performance of randomly selected Indiana third graders on the state mandated ISTEP+ standardized achievement test as a function of class size, pupil to teacher ratio, and the presence of a Prime Time instructional aide, using hierarchical linear modeling. The preliminary report and the AERA 2002 paper that summarizes the analyses are archived as ERIC document ED466679 (Lapsley, Daytner, Kelley, and Maxwell, 2002). An earlier background report from the same evaluation team, prior to the Notre Dame group joining the project, is archived as ERIC document ED455220.
Sampling. School corporations were drawn by stratified
cluster sampling with two rules: 25% of corporations from each
of the nine Indiana educational service regions, and at least
one urban corporation per region, with the remainder
proportionally allocated across geographic categories (urban,
suburban, town, rural). The achieved sample was 61
corporations (78% of the target), 163 schools, 573 classrooms
as counted in the manuscript (586 by the
paste(corp, school, class) rule used here), and 10,927
students (49.6% female; 85% Caucasian, 9.2% African
American, 3.2% Hispanic). 4,016 students were in classrooms
with a Prime Time instructional assistant (ptia == 1),
6,765 in classrooms without (ptia == 2); the file here
shows 4,021 and 6,789 plus 117 ptia == 3 (other assistant
listed), with the small differences reflecting cleaning rules
applied between the manuscript count and the final SPSS file.
Instruments. Third graders sit for the ISTEP+ (Indiana
Statewide Testing for Educational Progress) in September of
the school term. The ISTEP+ is published by CTB/McGraw-Hill
and includes language arts, reading, and mathematics
assessments. Normal Curve Equivalent (NCE) composite scores
for these domains and for the total are the
ncread / nclang / ncmath / nctotal columns and were the
criterion variables in the published HLM analyses. NCE scores
have a population mean of 50 and a standard deviation of
approximately 21.06, with percentiles 1, 50, and 99 mapping to
NCE scores of 1, 50, and 99. The
ge* family is the parallel Gates-MacGinitie battery; the
aa* family is the African American comparison NCE
(AANCE); the npa* family is the cognitive abilities
battery used as student level covariates in Finch, Bolin, and
Kelley (2019).
Nested data structure. The natural hierarchy is
student within classroom within school
within corporation within region. The derived
identifiers corp_id, school_id, and
class_id are pre-computed and safe to use as grouping
variables; the bare corp and class columns are
not unique by themselves. Class sizes range from 3 to
28 students (median 19); schools have 1 to 8 third grade
classrooms (median 3) and 11 to 166 students (median 65);
corporations have 15 to 756 students (median 117). Variance
decomposition for the published outcome
nctotal based on the three level random intercept null
model lmer(nctotal ~ 1 + (1 | corp_id/school_id)) gives:
between-corporation variance 16.29, between-school within
corporation variance 22.72, and within school residual variance
240.43, so that
ICC\(_{\mathrm{corp}}\) \(\approx\) 0.058,
ICC\(_{\mathrm{school|corp}}\) \(\approx\) 0.081, and
the combined cluster
ICC\(_{\mathrm{cluster}}\) \(\approx\) 0.140. These
nontrivial intraclass correlations are the methodological
reason multilevel modeling is preferred to ordinary least
squares regression for these data.
Level 1, 2, 3 model framework. For an outcome
\(Y_{ijk}\) on student \(i\) in classroom \(j\) in school
\(k\), with student level predictor \(X^{(1)}_{ijk}\),
classroom level predictor \(X^{(2)}_{jk}\), and school level
predictor \(X^{(3)}_k\), the published Lapsley et al. (2002)
family of HLM models has the equations
$$Y_{ijk} = \pi_{0jk} + \pi_{1jk} X^{(1)}_{ijk} + e_{ijk}
\quad \text{(Level 1)},$$
$$\pi_{0jk} = \beta_{00k} + \beta_{01k} X^{(2)}_{jk} +
r_{0jk},
\quad \pi_{1jk} = \beta_{10k} + r_{1jk}
\quad \text{(Level 2)},$$
$$\beta_{00k} = \gamma_{000} + \gamma_{001} X^{(3)}_k +
u_{00k},
\quad \beta_{01k} = \gamma_{010},
\quad \beta_{10k} = \gamma_{100}
\quad \text{(Level 3)},$$
with \(e_{ijk} \sim N(0, \sigma^2)\), \(r_{jk} \sim
N(0, \mathbf{T}_\pi)\), and \(u_{00k} \sim N(0, \tau_{00})\).
Substituting upward, the reduced form is
$$Y_{ijk} = \gamma_{000} + \gamma_{100} X^{(1)}_{ijk} +
\gamma_{010} X^{(2)}_{jk} + \gamma_{001} X^{(3)}_k +
u_{00k} + r_{0jk} + r_{1jk} X^{(1)}_{ijk} + e_{ijk},$$
which in lme4 translates to
lmer(Y ~ X1 + X2 + X3 + (1 + X1 | corp_id/school_id)).
The examples give concrete fits as commented code, which the help
page therefore does not run; uncomment them to fit them.
Suggested benchmark uses. The data set is intentionally rich enough to support a wide range of demonstrations and benchmarks, including:
Two-, three-, and four-level random intercept and random slope models with lme4, nlme, or glmmTMB.
Cross-level interaction modeling (e.g., race \(\times\) class size, ptia \(\times\) SES).
ICC, design effect, and cluster level sample size calculations.
Comparisons of unweighted vs. design weighted estimators for stratified cluster samples.
Bayesian multilevel modeling and prior sensitivity (brms, MCMCglmm, rstanarm).
Missing data demonstrations (the
ge*,nc*, andaa*columns have non-trivial missingness; seevapply(prime_time_achievement, function(x) sum(is.na(x)), integer(1))).Multilevel reliability, intraclass correlation, and measurement invariance demonstrations across schools, corporations, and the categorical predictors.
Privacy and identifiability. The student level rows
contain no names, addresses, or other personally identifiable
information. Demographic variables are age in months, gender,
and a six category race code; all other fields are test scores
or aggregated school / corporation statistics. The numeric
corp, school, and class identifiers are
the same administrative numbers used in the original Indiana
Department of Education public files for the 2000 to 2001 school
year; they could in principle be cross referenced to that
public information to identify specific schools or
corporations. No individual student can be identified from any
combination of variables in this file.
Missing data convention. The Indiana DOE source used
999 as the student level missing data code and 888
as the "not applicable" code for typmulti and
ptstatus. The build script converts both to NA
(the SPSS missingness ranges already do most of the recoding on
import). The retained SPSS variable label is available via
attr(prime_time_achievement$X, "label") on every variable
carried over from the SPSS file (all columns except the derived
recode classize).
References
Primary citation. Lapsley, D. K., Daytner, K. M., Kelley, K., and Maxwell, S. E. (2002). Teacher aides, class size and academic achievement: A preliminary evaluation of Indiana's Prime Time. ERIC document ED466679. https://eric.ed.gov/?id=ED466679.
Use as a multilevel modeling running example. Finch, W. H., Bolin, J. E., and Kelley, K. (2019). Multilevel modeling using R (2nd ed.). CRC Press. The 2nd edition (Finch, Bolin, and Kelley, 2019) is the edition that uses these data; later editions are not authored by Kelley and should not be cited for that use.
Background report from the same evaluation team. Lapsley, D. K., and Daytner, K. M. (2001). Indiana's class size reduction initiative: Teacher perspectives on training, implementation, and pedagogy. ERIC document ED455220. https://files.eric.ed.gov/fulltext/ED455220.pdf.
Indiana statutory context. Indiana General Assembly,
House Bill 1166 (2001).
https://archive.iga.in.gov/2001/bills/IN/IN1166.1.html.
Project STAR background and Education Week coverage.
Pate-Bain, H., and Achilles, C. M. (1986). Interesting
developments on class size. Phi Delta Kappan, 67,
662–665. See also Education Week, Research: Sizing up
small classes (February 7, 2001),
https://www.edweek.org/leadership/research-sizing-up-small-classes/2001/02.
Project STAR teacher aide null result that motivated the Prime Time evaluation. Finn, J. D., Gerber, S. B., Farber, S. L., and Achilles, C. M. (2000). Teacher aides: An alternative to small classes? In M. C. Wang and J. D. Finn (Eds.), How small classes help teachers do their best (pp. 131–174). Temple University Center for Research in Human Development and Education.
Examples
data(prime_time_achievement)
dim(prime_time_achievement)
#> [1] 10927 113
# Variable labels from the SPSS source are preserved on every column:
attr(prime_time_achievement$nctotal, "label")
#> [1] "NCE TOTAL"
attr(prime_time_achievement$ptia, "label")
#> [1] "PRESENCE OF A PRIME TIME IA?"
# Cluster counts (reconciled with Lapsley et al., 2002):
length(unique(prime_time_achievement$corp_id)) # 61
#> [1] 61
length(unique(prime_time_achievement$school_id)) # 163
#> [1] 163
length(unique(prime_time_achievement$class_id)) # 586
#> [1] 586
# Reconciling with the manuscript:
table(prime_time_achievement$gender, useNA = "ifany")
#>
#> 1 2 <NA>
#> 5425 5457 45
table(prime_time_achievement$race, useNA = "ifany")
#>
#> 1 2 3 4 5 6 <NA>
#> 16 995 62 348 9207 188 111
table(prime_time_achievement$ptia)
#>
#> 1 2 3
#> 4021 6789 117
table(prime_time_achievement$classize)
#>
#> 1 2 3 4
#> 1085 4571 4854 417
# ----- Selecting subsets of interest -----
# Caucasian and African American only (the matched-race
# supplementary analyses in Lapsley et al., 2002):
pt_wb <- subset(prime_time_achievement, race %in% c(2L, 5L))
# Drop the few "other assistant listed" cases for a clean
# aide / no-aide contrast:
pt_clean <- subset(prime_time_achievement, ptia %in% c(1L, 2L))
# Only rural corporations (geog == 4), which is what the source
# SPSS file's FILTER_$ variable encoded:
pt_rural <- subset(prime_time_achievement, geog == 4L)
# Complete cases on the nctotal-on-race-and-class-size analysis:
analysis_vars <- c("nctotal", "race", "classize", "ses",
"corp_id", "school_id", "class_id")
pt_complete <- prime_time_achievement[
complete.cases(prime_time_achievement[, analysis_vars]),
analysis_vars
]
# ----- Multilevel fits -----
# The three fits below are shown but not run, because each one
# estimates a multilevel model on the full student level file and
# together they cost more time than a help page should take.
# Uncomment to fit them.
# Three-level null random intercept model. The variance
# decomposition gives the corp, school | corp, and within
# ICCs reported in the Details section.
# m_null <- lme4::lmer(nctotal ~ 1 + (1 | corp_id/school_id),
# data = prime_time_achievement)
# summary(m_null)
# Level-1 (race), level-2 (ptia and classize), and level-3
# (ses) main-effect model. Compare to Lapsley et al. (2002),
# which fit closely related HLM specifications.
# m_main <- lme4::lmer(
# nctotal ~ factor(race) + factor(ptia) + classize + ses +
# (1 | corp_id/school_id),
# data = prime_time_achievement)
# summary(m_main)
# Cross-level interaction: ptia x ses (the published finding
# was that aide benefit was concentrated in higher-SES
# schools).
# m_inter <- lme4::lmer(
# nctotal ~ factor(race) + factor(ptia) * ses + classize +
# (1 | corp_id/school_id),
# data = prime_time_achievement)
# summary(m_inter)