Public Health Data Analysis: HIV Viral Load in Atlanta
This paper presents a two-part public health data analysis of HIV/AIDS prevalence among 359 cases in the Atlanta metropolitan area. Part One applies descriptive epidemiology to examine viral load distributions by place (city of residence), person characteristics (gender and ethnicity), and American Indian/Alaskan Native subgroup comparisons. Atlanta accounts for nearly half of all cases, while females and African Americans report disproportionately higher viral loads. Part Two applies analytical epidemiology, using linear regression to test whether age significantly predicts HIV viral load. The regression results, supported by Pearson's and Spearman's correlation analyses, indicate that age is not a statistically significant predictor of viral load in this dataset.
- Introduction and Dataset Overview: Dataset scope: 359 HIV/AIDS cases, multiple variables
- Descriptive Epidemiology: Place: Viral load distribution across Atlanta-area cities
- Descriptive Epidemiology: Person Characteristics: Gender and ethnicity differences in HIV viral load
- Analytical Epidemiology: Hypothesis and Method: Age-viral load hypothesis and linear regression rationale
- Linear Regression Results and Interpretation: Regression output shows age not significant predictor
- Conclusion: Age does not significantly predict HIV viral load
✍️ How to write this paper — guide, tools & examples ▾
What makes this paper effective
- Clearly separates descriptive and analytical epidemiology into distinct sections, making the logical progression from data description to hypothesis testing easy to follow.
- Justifies the choice of statistical test (linear regression over logistic regression) by explaining the nature of the outcome variable, demonstrating methodological awareness.
- Honestly reports a null finding — that age is not a significant predictor of viral load — without distorting the data to fit the hypothesis, reflecting good scientific integrity.
Key academic technique demonstrated
The paper demonstrates hypothesis-driven quantitative analysis in public health. By stating a directional hypothesis before running the regression and then evaluating it against p-values from multiple correlation tests (F-statistic, Pearson's, and Spearman's), the author applies a complete inferential workflow rather than merely reporting descriptive statistics.
Structure breakdown
The paper is organized in two clearly labeled parts. Part One covers descriptive epidemiology across three classic epidemiological dimensions — place, person, and a subgroup comparison — supported by frequency tables and charts generated in Epi Info 7. Part Two states a testable hypothesis, justifies the chosen statistical method, presents the full regression output (coefficients, confidence intervals, F-statistics, and correlation coefficients), and interprets the results against the original hypothesis. A brief reference list closes the paper.
Introduction and Dataset Overview
The dataset selected for analysis concerns HIV/AIDS. It presents the prevalence of HIV/AIDS among 359 cases across several variables, including gender, ethnicity, city of residence, state, age, and sexual orientation. The analysis is divided into two parts: descriptive epidemiology, which characterizes the distribution of viral load across place and person, and analytical epidemiology, which tests a hypothesis about the relationship between age and viral load using linear regression.
Descriptive Epidemiology: Place
The highest occurrence of HIV/AIDS is reported in the City of Atlanta, which accounts for 47.16 percent of HIV/AIDS cases among the 359 cases. The second-highest occurrence is reported in College Park at 8.93 percent, followed by Alpharetta at 7.27 percent. The lowest HIV/AIDS occurrence is reported in Hapeville and Johns Creek, both of which report a prevalence rate below 2 percent.
The frequency table below (Fig. 1) presents the frequency of viral load weighted by city of residence, based on 359 records analyzed using Epi Info 7.
Fig. 1: Viral Load by City of Residence
Descriptive Epidemiology: Person Characteristics
The viral load is higher among females at 51.62 percent compared to males, who report a viral load frequency of 48.38 percent. These findings are summarized in the frequency table below (Fig. 2).
Fig. 2: Frequency Table of Viral Load by Gender
Figure 3 summarizes the person characteristics of the dataset by ethnic grouping. African Americans report a higher viral load frequency compared to Asians and Alaskan Natives, despite the fact that whites constitute the largest percentage of the overall sample, as shown in the combined frequency table in Figure 4.
The means table indicates that the mean viral load among American Indian/Alaskan Natives was 5,280, compared to an average of 4,500 for non-American Indian/Alaskan Natives. Thus, American Indian/Alaskan Natives report a higher mean viral load than the general American population represented in this dataset.
Analytical Epidemiology: Hypothesis and Method
The hypothesis developed for this part of the analysis is:
Age significantly influences HIV viral load, with younger people reporting higher viral loads.
A linear regression was used to test this hypothesis. Linear regression is used to predict the relationship between variables and to assess the effect of one variable (the independent variable) on another (the dependent variable) (CDC, n.d.). The hypothesis focuses on determining the degree to which age influences HIV viral load. Linear regression is preferred over logistic regression because the outcome (dependent) variable — viral load — is a numerical, continuous variable (CDC, n.d.). Logistic regression is appropriate when the outcome variable is binary, taking on only two values, such as yes or no (CDC, n.d.). The continuous nature of the outcome variable therefore makes linear regression the most appropriate advanced statistical test (CDC, n.d.).
A complex sample means test could indicate which age categories have the highest viral loads based on calculated means per age group; however, it would not reveal the strength of the relationship between the two variables, making linear regression the more informative choice.
Conclusion
The regression results indicate that age is not a statistically significant predictor of HIV viral load in this dataset, as evidenced by p-values exceeding 0.05 across all correlation tests. Descriptively, the data reveal meaningful geographic and demographic disparities: Atlanta accounts for nearly half of all cases, females carry a slightly higher viral load than males, and American Indian/Alaskan Natives exhibit a higher mean viral load than the broader population. These epidemiological patterns underscore the importance of targeted public health interventions that account for both geographic concentration and demographic risk factors.
References
CDC (n.d.). Visual Dashboard: Performing Statistical Analysis with Visual Tools. Centers for Disease Control and Prevention. Retrieved from https://www.cdc.gov/epiinfo/pdfs/userguide/8_visualdashboard.pdf
Always verify citation format against your institution’s current style guide requirements.