Project Marketplace
Project Topics & Materials
Search curated materials across every major department, faculty and institution.
Showing 32 materials
RISK MODELING IN CRYPTOCURRENCY MARKETS
Admin
About This Research Topic Cryptocurrency markets encompassing Bitcoin and Ethereum have experienced rapid growth in trading volume and investor participation globally, with Nigeria consistently ranking among world's leading adoption markets reflecting currency depreciation concerns, remittance use cases, and speculative interest. This substantial participation persists notwithstanding evolving regulatory stance and well-documented extreme volatility. Risk modeling in cryptocurrency markets — statistical quantification of potential magnitude of financial loss — is particularly important in crypto markets given documented tendency towards extreme swings, fat-tailed return distributions, and pronounced volatility clustering, characteristics more pronounced than in traditional equity markets. Value-at-Risk (VaR) and Expected Shortfall (Conditional VaR) represent industry-standard risk metrics quantifying potential loss at specified confidence over specified horizon. Choice of methodology carries material consequences: simpler parametric approaches assuming normal returns are computationally straightforward but risk substantially understating true tail risk in fat-tailed crypto markets, while sophisticated GARCH-based approaches explicitly modeling time-varying volatility can provide more well-calibrated estimates at cost of complexity. This study applies and comparatively evaluates historical simulation, parametric normal, and GARCH-based VaR to Bitcoin and Ethereum daily returns benchmarked against NGX All-Share Index over three-year illustrative period, identifying best-calibrated approach and assessing diversification between crypto and traditional Nigerian equity exposure. Main Abstract Cryptocurrency markets have attracted substantial investor interest in Nigeria, one of the world's largest cryptocurrency adoption markets by several published rankings, notwithstanding regulatory ambiguity and pronounced price volatility characteristic of these emerging digital asset markets. This study statistically models the risk characteristics of cryptocurrency markets, focusing on Bitcoin and Ethereum daily price series benchmarked against the NGX All-Share Index, using illustrative data spanning a three-year historical period. Value-at-Risk (VaR) and Expected Shortfall (Conditional VaR) were estimated using historical simulation, variance-covariance (parametric normal), and GARCH-based approaches, with backtesting conducted via the Kupiec Proportion of Failures test. Descriptive statistics confirmed pronounced excess kurtosis and volatility substantially exceeding that of the NGX benchmark for both cryptocurrencies. GARCH(1,1) modelling confirmed strong volatility clustering in both Bitcoin (alpha + beta = 0.968) and Ethereum (alpha + beta = 0.951) return series, indicating high volatility persistence characteristic of speculative digital asset markets. The 99% one-day VaR, estimated via GARCH-based conditional volatility, was found to be statistically well-calibrated for Bitcoin (Kupiec test p = 0.412, failing to reject adequate calibration) but notably miscalibrated for the parametric normal approach for both assets (Kupiec test p < 0.05, rejecting adequate calibration), reflecting the parametric method's failure to account for fat-tailed return distributions. A Pearson correlation analysis revealed a statistically significant but modest positive correlation between Bitcoin and NGX All-Share Index returns (r = 0.184, p = 0.012), suggesting limited but non-trivial diversification benefit from combining cryptocurrency and traditional equity holdings. The study concludes that GARCH-based risk models substantially outperform simpler parametric approaches for cryptocurrency risk estimation given the pronounced volatility clustering and fat-tailed characteristics of these markets, and recommends that Nigerian investors and regulators adopt statistically robust, volatility-adaptive risk measurement approaches for cryptocurrency exposure. Keywords: cryptocurrency, Value-at-Risk, GARCH, risk modeling, Bitcoin, Ethereum, Expected Shortfall, Kupiec test, Nigeria
Statistical Analysis of Loan Default Factors
Admin
About This Research Topic Microfinance institutions occupy critical position within Nigerian financial sector, extending credit to individuals and micro-enterprises typically underserved by deposit money banks, supporting financial inclusion and micro-enterprise development. Sustainability of this function depends critically on effective credit risk management, with loan default as primary risk threatening institutional solvency and capacity to continue lending. At SCHOLARNESTHUB, we transform credit risk research into SEO-optimized academic articles. This study on statistical analysis of loan default factors is crafted for students searching for banking and finance project topics and statistics project topics . Loan default, defined as 90+ days past due, imposes direct costs through loss of principal and interest plus indirect costs via provisioning and constrained lending. Binary logistic regression provides standard tool modelling default probability as function of predictors, widely applied internationally and increasingly in Nigerian microfinance. Complementing binary approach, survival analysis including Cox proportional hazards model offers insight into not only whether loan defaults but when relative to origination — information relevant to portfolio monitoring and provisioning timing. This study applies both to 2,000 anonymised loan records with 14.6% default rate. Main Abstract Loan default remains significant operational and financial risk challenge for Nigerian microfinance institutions, with elevated rates threatening portfolio sustainability and constraining capacity to extend credit to underserved populations. This study statistically analyses determinants of loan default using anonymised sample of 2,000 loan records drawn from anonymised loan portfolio of selected Nigerian microfinance bank, applying binary logistic regression as primary technique complemented by chi-square tests and Cox proportional hazards survival analysis of time-to-default. Loan default operationalised as binary outcome (default vs non-default, 90+ days past due), with borrower demographic, loan characteristics and credit history variables examined as candidate predictors. Descriptive statistics revealed overall default rate 14.6% within sample. Binary logistic regression identified loan-to-income ratio (OR=2.68, p<0.001), prior default history (OR=3.84, p<0.001), collateral absence (OR=2.12, p<0.001), and loan tenor (OR=1.42, p=0.008) as significant positive predictors, while guarantor presence was significant negative predictor (OR=0.52, p=0.003). Model explained 38.9% variation (Nagelkerke R²=0.389) and correctly classified 84.2% cases. Cox proportional hazards analysis further confirmed loan-to-income ratio and prior default history as significant hazard-increasing covariates (p<0.001), with median survival time to default approximately 14 months among eventual defaulters. Study concludes loan-to-income ratio and prior default history are strongest statistical predictors of default risk, and recommends enhanced affordability assessment and credit history verification in loan origination.
CUSTOMER SATISFACTION ANALYSIS IN DIGITAL BANKING
Admin
About This Research Topic Nigerian banking sector has undergone substantial digital transformation over past decade with mobile apps internet banking and USSD increasingly displacing branch-based banking as primary interface for routine transactions. This shift has altered customer experience locus from in-branch interpersonal interactions toward digital interface design transaction reliability and remote support responsiveness. Customer satisfaction central construct in services marketing refers to overall evaluative judgment relative to expectations with downstream implications for loyalty retention and word-of-mouth referral critical in competitive digital landscape. SERVQUAL framework by Parasuraman Zeithaml and Berry 1988 provides most widely applied measurement decomposing service quality into five dimensions: reliability ability to perform promised service dependably accurately, responsiveness willingness to help and provide prompt service, assurance knowledge ability to inspire trust, empathy caring individualized attention, tangibles appearance of facilities and materials in digital context interface design. This article for SCHOLARNESTHUB presents rewritten SEO-optimized analysis of survey of 384 respondents determined via Taro Yamane formula using 25-item 5-point Likert scale across SERVQUAL dimensions modelling overall satisfaction via multiple regression. For similar service quality studies see banking and finance project topics on SCHOLARNESTHUB . Main Abstract Customer satisfaction remains critical determinant of competitive advantage and retention within Nigeria increasingly digitalised banking sector in which mobile applications internet banking platforms and USSD channels have become primary touchpoints. Study statistically analyses customer satisfaction with digital banking services among customers of selected Nigerian deposit money bank digital channels in selected city Nigeria applying SERVQUAL framework. Structured questionnaire incorporating 25-item 5-point Likert scale spanning five SERVQUAL dimensions reliability responsiveness assurance empathy tangibles administered to sample 384 respondents determined using Taro Yamane formula. Multiple linear regression employed to model overall satisfaction as function of five dimension scores complemented by one-way ANOVA for satisfaction differences across bank type and independent-samples t-tests for gender differences. Descriptive revealed mean overall satisfaction 3.72 SD 0.68 on 5-point scale. Multiple regression explaining 58.7 percent variance R2 0.587 F(5,378) 107.3 p<0.001 identified reliability beta 0.312 p<0.001 responsiveness beta 0.268 p<0.001 and assurance beta 0.184 p 0.002 as strongest predictors with empathy and tangibles also significant but smaller magnitude. One-way ANOVA revealed significant differences across bank type F(2,381) 8.94 p<0.001 while t-test found no significant gender difference t(382) 1.14 p 0.256. Study concludes reliability and responsiveness primary drivers of digital banking satisfaction and recommends prioritised investment in transaction reliability and support responsiveness.
STATISTICAL DETERMINANTS OF FINANCIAL INCLUSION
Admin
About This Research Topic Financial inclusion broadly defined as process of ensuring access to appropriate financial products and services needed by all segments of society in fair transparent equitable manner at affordable cost has been central pillar Nigerian economic development policy since launch Central Bank of Nigeria National Financial Inclusion Strategy in 2012. Despite sustained policy attention considerable expansion both traditional banking and digital financial service infrastructure successive Enhancing Financial Innovation and Access EFInA Access to Finance surveys documented persistent gaps between national financial inclusion targets and actual measured inclusion rates with substantial variation across demographic socioeconomic geographic segments. Understanding statistical determinants of financial inclusion that is specific individual and contextual characteristics that most strongly significantly predict formal financial access essential for designing effectively targeted policy interventions capable closing persistent inclusion gap. Binary logistic regression provides standard statistical tool for this purpose enabling researchers quantify independent statistical contribution multiple candidate determinants including education income geographic location technology access to probability formal financial inclusion while appropriately controlling simultaneous influence other correlated factors. Beyond individual-level determinant analysis financial inclusion itself frequently measured as multidimensional construct encompassing access proximity availability financial access points usage actual utilisation financial products services and quality extent available services meet users genuine needs. Principal Component Analysis offers statistically rigorous technique constructing composite Financial Inclusion Index from multiple underlying access usage indicators reducing dimensionality while preserving maximum proportion underlying variance methodological approach study applies complement primary binary logistic regression analysis. Recent analysis of EFInA 2023 nationally representative survey on 26361 observations covering demographic profiling barrier analysis hypothesis testing logistic regression and natural experiment Naira redesign policy shock and logistic regression analysis of women's financial inclusion using 2017 EFInA survey found education income urban residence mobile phone ownership strongest predictors. EFInA 2023 nationally representative survey 26361 observations logistic regression barrier analysis determinants women's financial inclusion Nigeria 2017 EFInA survey logistic regression education income urban residence mobile phone ownership This study applies both binary logistic regression and Principal Component Analysis to household survey data collected within study area with aim statistically identifying significant determinants financial inclusion and constructing robust composite inclusion index for supplementary comparative analysis. For related project materials see ScholarNestHub finance collection. ScholarNestHub finance collection Main Abstract Financial inclusion defined as availability and equality of access to useful and affordable financial products and services remains central pillar Nigeria economic development strategy yet substantial segments adult population continue lack access formal financial services. This study statistically examines determinants of financial inclusion among adults within SELECTED STATE/GEOPOLITICAL ZONE Nigeria drawing on structured household survey administered to 384 respondents determined using Taro Yamane formula. Financial inclusion operationalised as binary outcome formally included versus excluded and multiple binary logistic regression employed to model inclusion status as function demographic socioeconomic geographic predictors complemented by chi-square tests association and composite Financial Inclusion Index constructed via Principal Component Analysis. Descriptive statistics revealed overall financial inclusion rate 64.6% among sampled respondents. Binary logistic regression identified educational attainment OR=2.94 p<0.001 monthly income OR=2.18 p<0.001 proximity to financial access point OR=1.87 p=0.004 and mobile phone ownership OR=3.42 p<0.001 as statistically significant positive predictors financial inclusion while rural residence was statistically significant negative predictor OR=0.48 p=0.002. Model explained 47.6% variation inclusion status Nagelkerke R2=0.476. Principal Component Analysis reduced eleven financial access and usage indicators to three interpretable components jointly explaining 68.3% total variance providing robust composite Financial Inclusion Index used for supplementary regional comparison. Study concludes financial inclusion in study area significantly shaped by education income mobile phone access geographic proximity financial infrastructure and recommends targeted mobile-money-led inclusion strategies for rural lower-income population segments. Keywords: financial inclusion, logistic regression, principal component analysis, financial access, Nigeria.
PREDICTIVE ANALYSIS OF STOCK MARKET PERFORMANCE
Admin
About This Research Topic The Nigerian capital market, anchored by the Nigerian Exchange Group (NGX), serves as critical channel for capital formation and investment, where companies raise long-term capital and investors allocate savings. The NGX All-Share Index, primary benchmark, reflects aggregate listed equities and is watched as barometer of market and broader economic sentiment. Predictive analysis of stock market performance has long attracted academic and practitioner interest due to potential rewards for anticipating price movements. Statistical time series methods including Box-Jenkins ARIMA and GARCH-family volatility models provide rigorous toolkit for examining whether historical price and return patterns carry exploitable predictive information. This inquiry connects directly to Efficient Market Hypothesis (EMH) holding that prices fully reflect available information implying consistent prediction should not be possible. Testing return and volatility predictability therefore assesses degree of efficiency characterising Nigerian market - empirical question with substantial existing but unsettled literature. This study applies ARIMA and GARCH to illustrative NGX All-Share daily closing data spanning five-year period, examining both return predictability and volatility predictability, and examines statistical relationship between crude oil price movements and NGX returns given Nigeria's oil-dependent macroeconomic structure. Practical stakes extend beyond academic interest: pension administrators, insurers, institutional investors rely on statistically grounded risk assessment for regulatory capital adequacy, while individual investors benefit from improved understanding of market's genuine predictability characteristics. Main Abstract Accurate prediction of stock market performance remains a subject of enduring interest to investors, portfolio managers and financial regulators, offering the potential to inform investment decision-making and risk management practice within Nigeria's capital market. This study applies time series statistical models to predict the performance of the Nigerian Exchange Group (NGX) All-Share Index, using illustrative daily closing price data spanning a five-year historical period. Preliminary tests for stationarity using the Augmented Dickey-Fuller test confirmed that the raw price series is non-stationary, consistent with the weak-form Efficient Market Hypothesis, while the log-return series was found to be stationary. The Box-Jenkins ARIMA methodology was applied to the return series, with model identification guided by the Akaike Information Criterion and residual diagnostic checking, and further benchmarked against a GARCH(1,1) volatility model given evidence of volatility clustering identified through the ARCH-LM test. The selected ARIMA(1,0,1)-GARCH(1,1) model was validated using out-of-sample forecast evaluation. Results indicate that daily returns exhibit only weak, marginally significant autocorrelation (consistent with near-random-walk behaviour), while volatility exhibits strong, statistically significant clustering and persistence (GARCH parameters alpha + beta = 0.93, indicating high volatility persistence). Granger causality testing further revealed a statistically significant unidirectional relationship from crude oil price changes to NGX All-Share Index returns, consistent with Nigeria's oil-dependent macroeconomic structure. The study concludes that while short-term directional return prediction remains statistically limited, consistent with market efficiency, volatility forecasting using GARCH-family models offers a statistically robust and practically useful tool for risk management purposes within the Nigerian capital market. Keywords: stock market prediction, ARIMA, GARCH, Granger causality, market efficiency, Nigeria, NGX All-Share Index, volatility clustering
Fraud Detection in Digital Payment Systems
Admin
About This Research Topic The rapid digitalisation of Nigeria's payment ecosystem - mobile banking, internet banking, POS terminals and cards - driven by CBN cashless policy and fintech growth, has delivered convenience and inclusion but created new fraud vectors. Fraud imposes direct losses, reputational damage and compliance burden. At SCHOLARNESTHUB, we transform data science projects into SEO-optimized academic articles. This study on fraud detection in digital payment systems is tailored for students searching for computer science project topics and banking and finance project topics. Statistical and machine learning classification - logistic regression, decision tree and random forest - provides toolkit to distinguish fraudulent from legitimate transactions with accuracy unattainable via manual rule-based review. This article applies and compares three approaches on 5,000 illustrative transactions with 2.4% fraud rate, using transaction amount, time-of-day, channel, velocity and historical behaviour features, with NIBSS fraud landscape context. Main Abstract Rapid expansion of digital payment channels in Nigeria accompanied by rise in digital payment fraud posing financial and reputational risk. This study applies statistical and machine learning classification to anonymised sample of digital payment transactions from Nigerian deposit money bank to develop and evaluate robust fraud detection model. Dataset of 5,000 illustrative records incorporating amount, time-of-day, channel, velocity and historical behaviour features, class-imbalanced fraud incidence 2.4%, analysed using logistic regression, decision tree and random forest, benchmarked using classification metrics. Exploratory statistics revealed significant differences between fraudulent and legitimate transactions across key features confirmed via t-tests and chi-square. Random forest achieved strongest discriminatory performance (AUC-ROC=0.947) outperforming logistic regression (0.891) and decision tree (0.872), with velocity, amount deviation from historical average, and unusual transaction time as most important features. Precision-recall analysis appropriate given imbalance confirmed superior balance between sensitivity and false-positive rate. Study concludes ensemble methods informed by statistically validated behavioural features offer robust superior approach relative to simpler baselines and recommends integration into real-time monitoring alongside continuous performance monitoring.
Predicting Water Resource Availability Using Time Series Models
Admin
About This Research Topic Water is foundational — underpinning irrigation, hydropower, domestic supply and ecological stability. In Nigeria, Niger-Benue system supplies irrigation schemes, hydroelectric stations and municipal works for millions, yet availability is highly variable driven by bimodal rainfall, upstream abstraction, land use change and climate variability. At SCHOLARNESTHUB, we transform statistics and hydrology research into SEO-optimized academic articles. This study on predicting water resource availability using time series models is built for students searching for statistics project topics and environmental science project topics . Unlike descriptive summaries, ARIMA family explicitly captures trend, seasonality and autocorrelation. Box-Jenkins methodology has become standard in hydrological forecasting due to flexibility accommodating non-stationary seasonal series via differencing. Recent records suggest increasing volatility in wet-season peaks raising flood risk and dry-season minimums threatening supply during Harmattan, underscoring urgency of robust statistical forecasting. This study applies formal methodology to 20-year monthly streamflow and rainfall records from selected Lower Benue River Basin stations across Benue and Kogi States. Main Abstract Reliable prediction of water resource availability is central to effective planning of irrigation, hydropower, domestic supply and flood control, particularly in basins with considerable seasonal and inter-annual variability. This study applies time series models to monthly streamflow and rainfall records from Lower Benue River Basin covering selected gauging stations across Benue and Kogi States to model and forecast availability. Secondary monthly data spanning twenty-year period were obtained from illustrative hydrological records and subjected to preliminary tests for stationarity using Augmented Dickey-Fuller (ADF) and Phillips-Perron (PP) tests. Classical decomposition and Box-Jenkins ARIMA methodology were employed, with model identification guided by Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), and diagnostic checks on residual autocorrelation. Seasonal ARIMA (SARIMA) was found to outperform non-seasonal specifications given pronounced twelve-month periodicity associated with Nigeria's bimodal rainfall pattern. Selected SARIMA(1,1,1)(1,1,1)12 model was validated using out-of-sample forecast evaluation achieving Mean Absolute Percentage Error (MAPE) within acceptable bounds for hydrological forecasting. Results indicate statistically significant declining trend in dry-season minimum flows alongside increasing variability in wet-season peak flows, both significant at 5% level. Study concludes time series forecasting provides valuable early-warning and planning tool for water resource managers in basin and recommends institutionalisation of continuous hydrological monitoring and periodic model recalibration.
STATISTICAL ANALYSIS OF WASTE MANAGEMENT PRACTICES
Admin
About This Research Topic Municipal solid waste management has emerged as one of the most pressing urban environmental challenges in Nigeria, driven by rapid population growth, urbanisation, and changing consumption patterns outpacing formal collection capacity. Improperly managed waste blocks drainage and causes flooding, breeds disease vectors, and creates public health risks especially in densely populated low-income neighbourhoods. Statistical analysis of waste management practices provides rigorous evidence-based foundation: descriptive stats quantify volume and composition, while chi-square, logistic regression and ANOVA test relationships between household characteristics and disposal behaviour beyond anecdotal assessment. In many LGAs, responsibility is vested in state waste agencies or private contractors, yet formal coverage remains partial, particularly peri-urban and informal neighbourhoods. Where collection unavailable, households resort to open dumping, burning, or burial. Consequences are not evenly distributed: lower-income and peri-urban areas bear disproportionate burden, raising equity dimension illuminated by socioeconomic-focused analysis. Scale is considerable: rapid growth has outstripped fleet and disposal site capacity, leading to visible accumulation in public spaces and drainage channels, attracting attention from policymakers yet evidence base remains thin relative to investment contemplated. Beyond health and environment, effective management intersects with urban planning, climate mitigation through methane from landfills, and circular economy via recycling. Cities that transitioned to higher formal collection did so through infrastructure, tariff reform, and behaviour-change communication informed by household-level baseline data - precisely evidence this study generates. Main Abstract Rapid urbanisation and population growth in Nigerian cities have intensified the challenge of municipal solid waste management, with implications for public health, environmental quality and urban aesthetics. This study statistically analyses waste management practices within [SELECTED LOCAL GOVERNMENT AREA], [STATE], Nigeria, focusing on household waste generation patterns, disposal behaviour, and the socioeconomic determinants of waste management practices. A structured questionnaire was administered to a sample of 384 households determined using the Taro Yamane formula, eliciting responses on a 5-point Likert scale alongside categorical data on waste disposal methods. Descriptive statistics (frequencies, percentages, mean, standard deviation) were used to characterise waste generation and disposal patterns, while inferential statistics, including Chi-Square tests of independence, binary logistic regression, and one-way ANOVA, were used to test formulated hypotheses regarding the relationship between socioeconomic status, awareness level, and waste management practices. Results show that 62.3% of sampled households practised improper waste disposal (open dumping or burning), with a statistically significant association found between household income level and waste disposal method (chi-square = 28.47, df = 4, p < 0.001). Binary logistic regression identified educational attainment and access to formal waste collection services as statistically significant predictors of proper waste disposal behaviour. The study concludes that waste management practices in the study area are significantly shaped by socioeconomic and infrastructural factors, and recommends expanded formal waste collection coverage alongside targeted public health education campaigns. Keywords: waste management, municipal solid waste, logistic regression, chi-square test, Nigeria, ANOVA, improper disposal
CLIMATE VARIABILITY AND FOOD SECURITY: A STATISTICAL APPROACH
Admin
About This Research Topic Food security conventionally defined as state in which all people at all times have physical and economic access to sufficient safe and nutritious food to meet dietary needs remains pressing developmental challenge across sub-Saharan Africa and Nigeria in particular where agriculture continues serve as primary livelihood source substantial share rural population. Overwhelming reliance Nigerian agriculture on rain-fed rather than irrigated production systems renders agricultural output and by extension food security outcomes highly sensitive to climatic conditions particularly rainfall timing quantity distribution as well as temperature regimes affecting crop physiological processes. Climate variability referring to fluctuations in climatic conditions around long-term average patterns over periods ranging season to season year to year increasingly implicated in observed volatility Nigerian agricultural output. Unlike long-term climate change which describes gradual shift average climatic conditions over decades climate variability captures shorter-term fluctuations including delayed onset rains mid-season dry spells anomalous temperature episodes that directly disrupt cropping calendars can precipitate acute season-specific production shortfalls with immediate food security consequences. Statistical analysis provides essential evidence-based lens through which relationship between climate variability and food security outcomes can be rigorously examined moving beyond anecdotal purely descriptive accounts specific drought or flood episodes towards formal quantified characterisation strength direction statistical significance climate-agriculture relationships over extended historical record. Correlation regression analysis in particular allow researchers isolate statistical contribution specific climatic variables to variation agricultural output while controlling simultaneous influence multiple climatic factors. Study applies statistical techniques to secondary climatic agricultural production data for study area with aim quantifying statistical relationship between climate variability indicators food security outcomes generating evidence to inform climate-adaptive agricultural policy extension planning. Urgency inquiry underscored broader global context increasing climatic volatility associated anthropogenic climate change which climate science projects will further intensify rainfall variability temperature extremes across West Africa coming decades. Understanding historical statistical relationship between climate variability food security within Nigerian context provides essential empirical foundation for anticipating planning for these projected future changes and for designing agricultural systems policies greater climate resilience. Nigeria National Agricultural Technology and Innovation Policy successive national development plans repeatedly identified climate resilience as strategic priority agricultural sector yet translation strategic priority into operational statistically grounded planning tools at zonal or local government level remains uneven. This study contributes towards closing translation gap by demonstrating for specific illustrative study area how routinely collected climatic agricultural production data can be statistically analysed to yield directly actionable evidence for local agricultural planning extension advisory design. Previous investigations on effect of climatic variability on maize production Nigeria correlation multivariate regression and climate variability maize yield correlation regression r squared 0.333 rainfall temperature Nigeria investigated variability climate parameters food crop yields Nigeria using correlation multivariate regression findings revealed pineapple more sensitive 76.17% while maize groundnut more stable and significant moderate positive relationship between temperature maize yield linear regression r squared 0.333 variation explained. For related project materials see ScholarNestHub agriculture collection . Mian Abstract Climate variability manifested through fluctuating rainfall patterns rising temperatures and increasingly frequent extreme weather events poses significant threat to agricultural productivity and food security in Nigeria economy in which substantial share population depends on rain-fed agriculture for livelihood subsistence. This study statistically examines relationship between climate variability indicators and food security outcomes within SELECTED AGRICULTURAL ZONE STATE Nigeria using secondary time series data spanning twenty-year period drawn from illustrative meteorological agricultural production records. Descriptive statistics characterised trend variability rainfall temperature crop yield series while Pearson correlation analysis and multiple linear regression were employed to quantify statistical relationship between climatic variables annual rainfall mean temperature rainfall variability index and food security indicators cereal crop yield per capita food production index. One-way ANOVA further tested significant differences in crop yield across classified rainfall-adequacy years. Results reveal statistically significant positive correlation between annual rainfall and cereal yield r=0.68 p<0.001 and statistically significant negative correlation between temperature anomaly and yield r=-0.52 p<0.001. Multiple regression model with rainfall temperature anomaly rainfall variability as predictors explained 61.4% variation in cereal yield R2=0.614 F(3,16)=8.47 p=0.001 with rainfall and temperature anomaly emerging as statistically significant individual predictors. ANOVA confirmed significantly lower yields in drought-classified years compared to normal and above-normal rainfall years F(2,17)=11.36 p=0.001. Study concludes climate variability exerts statistically significant influence on food security outcomes in study area and recommends integration climate-smart agricultural practices early-warning statistical forecasting into agricultural extension planning. Keywords: climate variability, food security, correlation, regression, ANOVA, Nigeria.
STATISTICAL MODELING OF MOBILE BANKING ADOPTION
Admin
About This Research Topic Mobile banking defined as use of mobile telecommunications devices to conduct financial transactions has rapidly expanded across Nigeria driven by high mobile phone penetration expanding telecom infrastructure and regulatory support from Central Bank of Nigeria financial inclusion strategy. This growth offers pathway for previously unbanked and underbanked Nigerians particularly in areas with limited traditional banking infrastructure to access formal savings payment and credit services. Despite expansion adoption remains uneven across demographic and geographic segments with variation associated with age educational attainment income and prior exposure to digital financial technology. Understanding statistical determinants is essential to financial institutions seeking to expand customer base and policymakers pursuing national financial inclusion targets where mobile channel is identified as most cost-effective mechanism versus physical branch expansion. Technology Acceptance Model and Unified Theory of Acceptance and Use of Technology provide well-established frameworks positing perceived usefulness perceived ease of use social influence facilitating conditions and perceived risk jointly shape behavioural intention. Statistical modelling particularly binary logistic regression provides appropriate quantitative tool for testing these theorised relationships. This article for SCHOLARNESTHUB presents rewritten SEO-optimized analysis of survey of 384 respondents with 68.2 percent adoption rate modelling predictors of mobile banking adoption. For similar quantitative frameworks see fintech project topics on SCHOLARNESTHUB . Main Abstract Mobile banking has emerged as central pillar of Nigeria's financial inclusion agenda offering channel through which previously unbanked and underbanked populations can access formal financial services without reliance on traditional brick-and-mortar infrastructure. Study statistically models determinants of mobile banking adoption among residents of selected city/LGA Nigeria drawing on Technology Acceptance Model and Unified Theory of Acceptance and Use of Technology as guiding frameworks. Structured questionnaire administered to sample of 384 respondents determined using Taro Yamane formula eliciting Likert-scale responses on perceived usefulness perceived ease of use perceived risk social influence and facilitating conditions alongside binary adoption outcome. Descriptive statistics characterised demographic and adoption profiles while binary logistic regression employed to model probability of adoption as function of theorised predictors and Chi-Square tests assessed association between adoption and categorical demographic variables. Results show overall adoption rate 68.2 percent among sampled respondents. Logistic regression model explaining 41.3 percent variation Nagelkerke R2 0.413 identified perceived usefulness OR 2.84 p<0.001 perceived ease of use OR 1.92 p 0.003 and social influence OR 1.68 p 0.011 as statistically significant positive predictors while perceived risk significant negative predictor OR 0.54 p 0.002. Chi-square analysis revealed statistically significant association between educational attainment and adoption status chi-square 24.61 p<0.001. Study concludes adoption significantly shaped by technology acceptance constructs consistent with established models and recommends targeted usability improvements and risk-communication strategies to accelerate adoption among underserved segments.
PREDICTIVE ANALYTICS FOR CUSTOMER CHURN USING LOGISTIC REGRESSION AND RANDOM FOREST
Admin
About This Research Topic Customer churn defined as discontinuation of customer relationship with service provider within defined observation period represents one of most persistent financially consequential challenges facing subscription-based industries including telecommunications banking insurance streaming media services. Strategic importance churn management well established marketing customer relationship management literature which long documented cost acquiring new customer typically substantially exceeds cost retaining existing one making accurate statistically grounded churn prediction high-value analytical capability for any subscription-based business. Proliferation customer relationship management CRM systems digital service platforms generated increasingly rich granular customer-level data encompassing contractual details billing history service usage patterns customer service interaction logs that collectively constitute rich substrate for statistical churn prediction modelling. Within analytical landscape binary logistic regression and random forest emerged as two most widely applied directly comparable churn prediction methodologies: logistic regression classical statistical technique offering directly interpretable coefficients odds ratios of considerable value for business stakeholder communication regulatory transparency and random forest ensemble machine learning technique capable capturing complex non-linear interactions among predictors without requiring analyst to pre-specify functional form. While numerous prior studies compared logistic regression against random forest for churn prediction substantial portion comparative literature relies on narrow evaluation methodology most commonly single train-test split evaluated via accuracy or limited subset classification metrics without incorporating fuller battery statistical validation techniques increasingly regarded as best practice rigorous applied predictive modelling research. This narrower approach risks both overstating reliability any single-split performance estimate given absence cross-validation to assess estimate stability and understating practical business relevance model comparison given absence explicit linkage between statistical classification performance and actual monetary costs benefits retention decision-making that any deployed churn model would ultimately inform. This study accordingly undertakes substantially more comprehensive statistical evaluation extending beyond simple accuracy comparison to incorporate five complementary evaluation dimensions: standard classification metrics accuracy precision recall F1-score AUC on held-out test set; stratified k-fold cross-validation to assess stability generalisability; McNemar test providing formal statistical significance testing paired classification agreement; probabilistic calibration assessment via Brier score evaluating reliability underlying predicted probabilities property direct importance for any application such as targeted retention campaign budgeting that relies on probability estimates rather than binary classifications alone; and cost-sensitive profit curve analysis explicitly incorporating asymmetric business costs retention intervention cost false-positive retention offer extended to customer who would not in fact have churned against benefit successful retention value true-positive customer correctly identified retained translating abstract classification performance into directly interpretable business-value metric. Comprehensive multi-dimensional framework directly addresses well-recognised gap between academic churn prediction benchmarking practice which frequently emphasises accuracy or AUC in isolation and fuller statistical business rigour genuinely required for responsible defensible churn model selection deployment. By triangulating findings across standard classification metrics cross-validation stability formal paired significance testing probabilistic calibration and cost-sensitive business-value translation study positioned to reveal genuinely nuanced trade-offs such as finding two models trade off precision accuracy against recall discrimination that single-metric comparison would entirely obscure providing decision-makers fuller multi-dimensional evidentiary basis genuinely required defensible selection. Recent reproducible workflows on customer churn calibrated probability 5-fold cross-validation AUC Brier score and logistic regression vs LightGBM ROC AUC PR AUC Brier score profit curve optimal threshold demonstrate fully reproducible workflow transforms raw data into calibrated predictions 5-fold CV AUC Brier reliability curve and profit curve optimal tau maximizing net retention profit. For related project materials see ScholarNestHub data science collection . Main Abstract Customer churn discontinuation customer relationship with service provider represents persistent threat to revenue stability long-term profitability in subscription-based industries motivating substantial academic industry interest in statistically robust churn prediction methodologies. This study undertook comprehensive statistical evaluation of Logistic Regression and Random Forest as competing approaches to customer churn prediction extending beyond simple accuracy comparison to incorporate stratified k-fold cross-validation formal paired significance testing McNemar test probabilistic calibration assessment Brier score and cost-sensitive profit-curve analysis explicitly incorporating asymmetric business costs retention intervention using dataset 1,500 telecommunications customer records comprising contractual billing service-usage demographic variables. Specific objectives were to determine overall churn rate describe customer characteristics examine bivariate relationships between candidate predictors and churn identify statistically significant predictors using binary logistic regression and rigorously compare Logistic Regression against Random Forest across accuracy-based discrimination-based calibration-based statistical-significance-based and business-value-based evaluation criteria. Descriptive statistics Pearson correlation independent samples t-tests Chi-square tests one-way ANOVA binary logistic regression 5-fold stratified cross-validation McNemar test Brier score calibration analysis and cost-sensitive profit curve analysis were employed. Results showed overall churn rate 31.67% 475 of 1,500 customers. Contract type χ2=108.03 p<0.001 and technical support subscription χ2=28.43 p<0.001 were both strongly associated with churn and churned customers had significantly shorter tenure 28.45 versus 35.40 months t=-8.325 p<0.001 significantly higher monthly charges 68.12 versus 63.70 t=3.916 p<0.001 and significantly more customer service calls 1.91 versus 1.50 t=5.634 p<0.001 than retained customers. Binary logistic regression identified tenure monthly charges customer service calls contract type technical support online security senior citizen status as statistically significant predictors McFadden pseudo R2=0.166. On held-out test set Random Forest achieved marginally higher accuracy 71.33% versus 70.89% and precision 55.30% versus 53.37% alongside modestly better lower Brier calibration score 0.1903 versus 0.1967 while Logistic Regression achieved substantially higher recall 66.43% versus 51.05% higher AUC 75.43% versus 73.62% and higher 5-fold cross-validated mean AUC 76.30% versus 75.31% with lower cross-validation variance. McNemar test found no statistically significant difference in two models paired classification error patterns χ2=0.016 p=0.901 indicating despite differing performance profiles across individual metrics neither model significantly outperforms other in overall paired classification agreement with ground truth. Cost-sensitive profit curve analysis incorporating assumed retention-offer economics projected substantially higher expected retention profit under Logistic Regression model ₦23,057.41 than under Random Forest ₦17,885.80 driven primarily by Logistic Regression superior recall and consequently greater capture at-risk customers eligible for retention intervention. Study concludes model selection between Logistic Regression and Random Forest for churn prediction should be explicitly grounded in specific business decision context rather than single default metric with Logistic Regression superior recall discrimination and under assumed cost structure superior projected retention profit making it stronger candidate proactive retention campaign targeting while Random Forest superior precision calibration may better suit applications prioritising minimisation unnecessary retention-offer costs. Recommended telecommunications providers adopt cost-sensitive profit-curve-based model evaluation rather than accuracy alone when selecting churn prediction models for deployment.
ANALYSIS OF CARBON EMISSION TRENDS AND ECONOMIC GROWTH IN NIGERIA
Admin
About This Research Topic Climate change driven by anthropogenic greenhouse gases represents defining environmental and developmental challenge of twenty-first century. Carbon dioxide primary GHG by volume released through fossil fuel combustion, industrial processes, flaring and biomass use has risen from 280 ppm pre-industrial to over 421 ppm in 2023 driving 1.1C warming, sea level rise and extreme weather. Understanding relationship between economic development and environmental quality has been central question in environmental economics for three decades. Environmental Kuznets Curve hypothesis proposed by Grossman and Krueger 1991 and named after Simon Kuznets posits inverted U-shaped relationship between per capita income and degradation: pollution initially rises as low-income countries prioritise production, then declines as incomes rise shifting preferences toward environmental quality and enabling cleaner technology. Nigeria as Africa's largest economy and most populous nation is major and growing emitter estimated at 120-140 million tonnes annually, third or fourth largest in Africa, driven by petroleum production and gas flaring, ageing transport fleet, industrial process emissions and near-universal biomass cooking and diesel generators. Economic growth has been volatile with strong growth 2003-2014 averaging above 7 percent interspersed with oil-price recessions, providing multi-decade variation to test EKC. This article for SCHOLARNESTHUB presents rewritten SEO-optimized analysis of 43-year time series 1980-2022 testing EKC for Nigeria using ARDL bounds testing framework. Students exploring similar econometric designs can see environmental economics project topics on SCHOLARNESTHUB . Main Abstract Nigeria as Africa's largest economy and most populous nation faces challenge of sustaining rapid economic growth while managing greenhouse gas emissions substantial and growing due to dependence on petroleum production and combustion flared gas transport emissions and large-scale biomass energy consumption. Environmental Kuznets Curve hypothesis posits inverted U-shaped relationship between per capita income and environmental degradation providing framework whether growth can eventually reduce emissions. Study tested EKC hypothesis for Nigeria using annual time series 1980-2022 43 years on CO2 emissions per capita GDP per capita energy consumption trade openness urbanisation rate and industrial value added. Autoregressive Distributed Lag bounds testing examined long-run co-integration and Error Correction Model estimated short-run adjustment. Study also applied Mann-Kendall trend test Augmented Dickey-Fuller unit root tests and Granger causality. ADF confirmed all variables I(1). ARDL bounds confirmed long-run co-integration F-statistic 6.847 exceeding 1 percent upper critical bound 4.26. Long-run ARDL estimates confirmed inverted U-shaped EKC: GDP per capita positive beta 0.847 p<0.001 and GDP per capita squared negative beta -0.0000412 p<0.001 turning point approximately USD 4,287 per capita 2015 constant prices. Nigeria current per capita income approximately USD 2,100 indicating still on ascending portion and CO2 expected to continue rising toward turning point. Energy consumption strongest positive predictor beta 0.612 p<0.001. Granger causality confirmed bidirectional causality between energy consumption and economic growth consistent with feedback hypothesis. ECM coefficient -0.387 p<0.001 indicates 38.7 percent of short-run deviation corrected within one year. All four null hypotheses rejected. Study recommends accelerating renewable transition solar wind hydro to decouple growth from emissions implementing carbon pricing through Nigeria Emission Trading Scheme expanding natural gas for cooking to replace biomass and pursuing energy efficiency standards for industrial and transport sectors.
STATISTICAL EVALUATION OF RENEWABLE ENERGY ADOPTION AMONG HOUSEHOLDS IN KWARA STATE, NIGERIA
Admin
About This Research Topic Nigeria faces one of the world's most severe electricity access crises. With grid access at only about 60% of the population and those connected experiencing 16 to 20 hours of daily outages, households and small businesses bear huge costs on diesel and kerosene. In Kwara State, home to 3.5 million people, KEDC serves approximately 285,000 metered customers but reliable supply reaches far fewer. This crisis has made off-grid renewable energy not just a climate solution but a daily necessity. renewable energy adoption among households — particularly solar PV lanterns, solar home systems, and solar mini-grids — is now a critical market and policy priority. Globally, solar module costs fell 89% from 2010 to 2022 per IRENA, while pay-as-you-go models from ENGIE, d.light, and Greenlight Planet have removed upfront cost barriers. Nigeria's Rural Electrification Agency and National Renewable Energy and Energy Efficiency Policy target 30% renewable in the mix by 2030, with solar as primary platform. Kwara State, with 5.5-6.0 kWh/m2/day irradiance, is technically ideal, yet adoption remains low and unequal. This study provides a comprehensive statistical evaluation of renewable energy adoption among 384 households across Ilorin South (urban), Asa (peri-urban), and Baruten (rural), using chi-square, binary and ordinal logistic regression, and willingness-to-pay contingent valuation to identify determinants and estimate affordability for evidence-based policy. Main Abstract Nigeria faces a severe electricity access crisis, with grid electricity reaching only approximately 60% of the population and those with access experiencing frequent outages averaging 16 to 20 hours per day in many states. Kwara State, with a population of approximately 3.5 million, reflects this national crisis: KEDC distribution company serves approximately 285,000 metered customers but supplies reliable electricity to a much smaller fraction. In this context, household adoption of off-grid renewable energy systems, particularly solar PV lanterns, solar home systems, and solar mini-grid connections, represents both a growing market and a critical policy priority. This study conducted a comprehensive statistical evaluation of renewable energy adoption among 384 sampled households in three LGAs of Kwara State: Ilorin South (urban), Asa (peri-urban), and Baruten (rural). The study applied descriptive statistics, chi-square tests of association, binary logistic regression, ordinal logistic regression, and willingness-to-pay (WTP) contingent valuation to identify the socioeconomic, attitudinal, and infrastructure determinants of household renewable energy adoption and to estimate the premium households are willing to pay for reliable clean energy. The adoption rate of any renewable energy technology was 47.4% overall, with significant LGA variation: 64.2% in Ilorin South, 44.5% in Asa, and 23.4% in Baruten. Solar PV lanterns were the most common adopted technology (28.4%), followed by solar home systems (12.5%), and solar mini-grid connection (6.5%). Binary logistic regression identified monthly household income (aOR = 3.247 per income category, p < 0.001), education level (aOR = 2.184 per level, p < 0.001), prior experience with grid outages (aOR = 1.987, p = 0.002), awareness of government solar programmes (aOR = 2.841, p < 0.001), and distance from nearest town (aOR = 0.624, p < 0.001) as significant independent predictors. Gender of household head was not significant (p = 0.487). Mean WTP for reliable solar electricity was N3,847 per month (95% CI: N3,612 to N4,082). All four null hypotheses were rejected. The study recommends targeted solar subsidy programmes for the lowest-income quintile, expansion of rural mini-grid deployment in Baruten and other rural LGAs, integration of renewable energy awareness into agricultural extension services, and development of a local solar technician training programme to address maintenance barriers to sustained adoption. Keywords: Renewable Energy, Solar PV, Technology Adoption, Logistic Regression, Willingness to Pay, Energy Access, Household Survey, Kwara State, Off-Grid Electrification
Modeling Deforestation Patterns Using Spatial Statistics in Cross River State, Nigeria
Admin
About This Research Topic Forest loss rarely happens evenly across a landscape. It creeps outward from roads, spreads along settlement edges, and clusters wherever access meets demand for farmland or timber. Knowing that deforestation clusters is one thing; knowing exactly where those clusters sit, and which factors predict them with statistical confidence, is what actually lets a forestry commission decide where to send patrols and where to build a buffer zone. That is the gap this study set out to close for Cross River State, home to roughly 40% of Nigeria's remaining tropical rainforest. This article rewrites and expands a research study applying spatial statistical methods, Moran's I, hotspot analysis, and spatial regression, to satellite-derived forest loss data for Cross River State between 2010 and 2023, in order to map where deforestation is concentrated and identify its strongest predictors. It sits alongside other applied statistics and environmental research in ScholarNestHub's project topics library , including a related study on air quality prediction using statistical and machine learning models in Lagos State . The sections below walk through the study's background, problem, objectives, and scope, before closing with answers to the questions most commonly asked about spatial statistics and deforestation modelling. Main Abstract Cross River State contains approximately 40% of Nigeria's remaining tropical rainforest, making it the most important remaining forest ecosystem in the country and one of the most biodiverse terrestrial habitats in Africa. Despite legal protections including the Cross River National Park and numerous forest reserves, deforestation continues at alarming rates driven by agricultural expansion, timber extraction, charcoal production, and infrastructure development. Understanding the spatial patterns and statistical drivers of deforestation is essential for designing effective, geographically targeted conservation interventions. This study applied spatial statistical methods to model deforestation patterns in Cross River State using Global Forest Watch forest loss data and satellite-derived land cover classification for the period 2010 to 2023. The analytical framework integrated spatial autocorrelation analysis (Global and Local Moran's I), hotspot analysis (Getis-Ord Gi*), spatial regression modelling (Spatial Lag Model and Spatial Error Model), binary logistic regression with spatial random effects, and descriptive spatial trend analysis. Cross River State lost 187,400 hectares of forest cover between 2010 and 2023, representing 18.7% of its 2010 forest extent of approximately 1,000,000 hectares. Annual forest loss accelerated from a mean of 10,800 hectares per year in 2010 to 2015 to 16,200 hectares per year in 2018 to 2023. Global Moran's I for forest loss rates confirmed significant positive spatial autocorrelation (I = 0.412, p < 0.001), indicating that deforestation clusters geographically rather than occurring randomly. Local Moran's I identified three primary high-high deforestation hotspot clusters: the Boki-Obudu border area, the Obanliku-Bekwarra axis, and the Abi-Yakurr western transitional zone. Spatial lag regression identified distance from the nearest road (B = -0.847, p < 0.001), distance from the nearest settlement (B = -0.412, p < 0.001), and LGA-level population density (B = 0.384, p < 0.001) as the three strongest spatial predictors of forest loss rates after controlling for spatial autocorrelation. The Cross River National Park boundary showed a significant protective effect (B = -2.147 for cells within the park, p < 0.001), and all four null hypotheses were rejected. The study recommends strengthening enforcement of the National Park boundary particularly in the Boki-Obudu hotspot, establishing road access buffers that restrict new agricultural clearing within 5 km of unpaved forest roads, and implementing community forestry programmes in the western transitional zone as alternatives to slash-and-burn agriculture.
Environmental Pollution and Health Outcomes in Delta State: A Statistical Study of Gas Flaring, Water Quality, and Community Health
Admin
About This Research Topic In parts of Delta State, a gas flare has been burning within sight of people's homes for longer than some residents have been alive. Everyone living nearby has a story about a persistent cough, a child's asthma, a relative's skin condition, but stories are not statistics, and policy decisions about where to enforce, where to intervene, and where to draw a buffer zone need numbers, not anecdotes. That gap, between the well-documented reality of pollution in the Niger Delta and the comparatively thin statistical evidence connecting specific pollutants to specific health outcomes, is what this study set out to close. This article rewrites and expands a research study statistically examining the relationship between environmental pollution indicators, gas flaring proximity, air pollutants, and water contamination, and health outcomes across three Local Government Areas in Delta State, Nigeria. It sits alongside other applied statistics and environmental research in ScholarNestHub's project topics library , including a related study on air quality prediction using statistical and machine learning models in Lagos State . The sections below walk through the study's background, problem, objectives, and scope, before closing with answers to the questions most commonly asked about pollution-health research in petroleum-producing communities. Main Abstract This study statistically assesses the associations between environmental pollution indicators and health outcomes across three Local Government Areas in Delta State, Nigeria: Warri South, a high petroleum-activity area; Ughelli North, an area of moderate petroleum activity; and Sapele, an industrial and urban area without direct petroleum extraction. The study responds to a well-documented gap in the literature: while the environmental and social costs of oil extraction in the Niger Delta are extensively described in qualitative and descriptive research, rigorous statistical analyses that combine objectively measured pollutant concentrations with health outcome data, and that quantify pollution-health associations through regression coefficients and odds ratios rather than perception surveys alone, remain rare for Delta State specifically. The study draws on environmental monitoring data spanning 2019 to 2023 alongside primary survey data collected from 384 sampled households between February and April 2024. It applies bivariate correlation analysis, multiple linear regression to model annual respiratory symptom frequency, and binary logistic regression to identify independent predictors of chronic respiratory disease diagnosis, while controlling for sociodemographic confounders. Environmental indicators examined include ambient sulphur dioxide and PM2.5 concentrations, proximity to active gas flare sites, and borehole water Total Dissolved Solids levels, set against health outcomes spanning respiratory symptoms, dermatological complaints, and physician-diagnosed chronic disease. The study is designed to produce, for the first time, a statistically grounded quantification of pollution-health associations specific to these Delta State communities, evidence intended to help the Delta State Ministry of Environment and the Delta State Ministry of Health prioritise enforcement action and health interventions according to measured health impact rather than general pollution presence. It further aims to contribute to the broader global literature on the health effects of gas flaring, an area where rigorous quantitative evidence remains limited relative to the global scale of the practice.
Air Quality Prediction Using Statistical and Machine Learning Models in Lagos State, Nigeria
Admin
About This Research Topic Lagos traffic is famous for the wrong reasons, but sitting in it does more than waste time. The exhaust, the idling generators, the dust that rolls in every harmattan season, all of it adds up to some of the most polluted air in West Africa, and most residents have no way of knowing on any given morning whether that day's air is merely bad or genuinely dangerous. Lagos currently has no system that tells people in advance. It reports what the air was like, not what it's about to become. This article rewrites and expands a research study that builds toward exactly that missing capability, comparing four statistical and machine learning models, Multiple Linear Regression, Random Forest, Gradient Boosting, and an LSTM neural network, to see which one best predicts next-day PM2.5 concentrations in Lagos State using three years of monitoring data. It sits alongside other applied machine learning work in ScholarNestHub's project topics library , including a related study on an explainable AI framework for medical diagnosis decision support . The sections below walk through the study's background, problem, objectives, and scope, before closing with answers to the questions most commonly asked about air quality prediction and PM2.5 modelling. Main Abstract Air pollution is a leading environmental health risk globally and in Nigeria, with urban centres like Lagos State experiencing chronically elevated concentrations of particulate matter (PM2.5 and PM10), nitrogen dioxide, sulphur dioxide, carbon monoxide, and ground-level ozone. Accurate prediction of air quality concentrations enables early warning systems, health risk communication, and targeted pollution control interventions. This study developed and compared four air quality prediction models for PM2.5 concentration in Lagos State using three years of daily air quality and meteorological data, January 2021 to December 2023, comprising 1,095 daily observations. The models compared were Multiple Linear Regression (MLR), Random Forest (RF), Gradient Boosting Machine (GBM), and a Long Short-Term Memory (LSTM) neural network, using predictor variables that included meteorological factors and anthropogenic emission proxies such as traffic density and industrial activity indices. Descriptive analysis revealed that Lagos State's mean PM2.5 concentration over the study period was 47.3 micrograms per cubic metre (SD = 28.4), substantially exceeding the WHO annual guideline of 5 micrograms per cubic metre and the Nigerian NESREA 24-hour standard of 35 micrograms per cubic metre. Harmattan season (November to February) concentrations were significantly higher, at a mean of 71.8 micrograms per cubic metre, than rainy season (May to September) concentrations, at a mean of 28.4 micrograms per cubic metre. On the test dataset, the final 20 percent of chronological data comprising 219 days, the GBM model achieved the best predictive performance: RMSE = 8.74 micrograms per cubic metre, MAE = 6.12, R-squared = 0.877. Random Forest was second (RMSE = 9.41, R-squared = 0.854), LSTM was third (RMSE = 10.23, R-squared = 0.831), and multiple linear regression performed worst but still adequately (RMSE = 13.87, R-squared = 0.741). GBM feature importance identified relative humidity (24.3%), wind speed (18.7%), month as a harmattan indicator (14.2%), and temperature (12.8%) as the four most important predictors, and all four null hypotheses were rejected. The study recommends deploying the GBM model as the operational air quality prediction tool in Lagos State's early warning system, expanding the monitoring station network from the current eight stations to at least twenty stations across all local government areas, and implementing targeted emission control measures during identified high-risk periods.
Statistical Analysis of ChatGPT Adoption Among University Students
Admin
About This Research Topic Artificial intelligence has rapidly moved into everyday academic life, and ChatGPT, launched by OpenAI in November 2022, became the fastest-growing consumer application in history — one million users in days, over 100 million in two months. Built on transformer-based generative architecture, it produces coherent text for essay writing, code generation, problem solving, translation and tutoring, positioning itself as always-available academic assistant while raising concerns about academic dishonesty, accuracy and critical thinking erosion. Technology Acceptance Model (TAM) by Davis posits perceived usefulness and perceived ease of use as primary determinants of intention to use new technology, extended in TAM2 and UTAUT to include social influence, facilitating conditions and experience. In Nigeria, internet penetration reached 55.4% in 2023 per Nigerian Communications Commission, with campuses as high digital activity clusters. Yet patterns, motivations and consequences of ChatGPT adoption among Nigerian undergraduates remain underexplored, leaving policy formation in empirical vacuum risking overly restrictive or insufficiently firm responses. This study conducts comprehensive statistical analysis among 200 undergraduates at University of Lagos using structured 25-item Likert questionnaire. Methods include descriptive statistics, Pearson correlation, multiple linear regression, one-way ANOVA and chi-square tests. Findings show 96.5% ever used ChatGPT, 66.5% weekly or more frequent, with perceived usefulness beta 0.287, ease of use 0.214, frequency 0.172 as strongest predictors of adoption intention, model explaining 68.3% variance. Significant differences across academic levels F=8.74 p<0.001 but no gender difference chi-square 0.184 p=0.912. The analysis demonstrates multivariate pipeline applicable to emerging AI phenomena in resource-constrained contexts. For methodological foundation, see our guides to technology adoption models and survey analysis using SPSS. Main Abstract This study undertook statistical analysis of ChatGPT adoption among university students focusing on drivers and perceived academic impact. Cross-sectional survey design with 200 undergraduate students at University of Lagos using structured 25-item Likert questionnaire. Data analysed via descriptive statistics, Pearson correlation, multiple linear regression, one-way ANOVA and chi-square tests. Majority 96.5% had used ChatGPT at least once, 66.5% reporting multiple times per week or more. Perceived usefulness (beta=0.287, p<0.001), ease of use (beta=0.214, p<0.001) and frequency of use (beta=0.172, p=0.006) emerged as strongest predictors of adoption intention. Significant difference in adoption levels across academic levels (F=8.74, p<0.001), while no significant gender difference (chi-square=0.184, p=0.912). Overall regression model accounted for 68.3% variance (R2=0.683, F=52.47, p<0.001). Study concludes ChatGPT adoption widespread principally driven by perceived utility and accessibility. Recommendations for universities to develop clear AI policies, integrate AI literacy into curricula and promote ethical use.
Statistical Methods for Detecting Fake News on Social Media
Admin
About This Research Topic Social media as primary news source has created vast, instantaneous and largely unregulated information ecosystem where fabricated content spreads faster than verified reporting. False stories on Twitter spread six times faster than true stories and reach far more users, amplifying risks to public health, elections and communal cohesion. In Nigeria, with over 109 million active internet users, false reports about ethnic violence and disease outbreaks have triggered real-world harm. Manual fact-checking cannot scale to 500 million tweets per day and billions of Facebook shares monthly. Automated detection is essential, yet many proprietary systems are opaque and models trained on Western datasets generalize poorly to multilingual, code-switched, WhatsApp-heavy Nigerian context. This study investigates transparent statistical methods — descriptive statistics, chi-square tests, binary logistic regression and TF-IDF text feature analysis — applied to 300 news articles sampled from Twitter and Facebook over six months. Results show number of shares (beta 0.412, p<0.001), source credibility score (beta -0.538, p<0.001) and emotional language presence (beta 0.319, p=0.003) significantly predict fake news classification, with overall accuracy 87.3%, sensitivity 84.6% and specificity 89.1%. Chi-square revealed significant association between platform type and fake news likelihood (X2 24.17, df 3, p<0.001). Logistic regression combined with NLP offers robust interpretable framework for resource-constrained environments. For practical implementation, see our guides to logistic regression and text mining with TF-IDF for social media data. Main Abstract Proliferation of fake news on social media threatens public discourse, democratic institutions and individual decision-making. This study investigates statistical methods for detecting fake news using 300 news articles sampled from Twitter and Facebook over six-month period. Descriptive statistics, chi-square tests of independence, binary logistic regression and text-based feature analysis including TF-IDF were employed. Logistic regression model indicated number of shares (beta=0.412, p<0.001), source credibility score (beta=-0.538, p<0.001) and presence of emotional language (beta=0.319, p=0.003) are significant predictors of fake news classification. Overall accuracy 87.3%, sensitivity 84.6%, specificity 89.1%. Chi-square revealed significant association between platform category and likelihood of fake news spread (X2=24.17, df=3, p<0.001). Study concludes statistical machine learning hybrids, particularly logistic regression combined with NLP feature extraction, offer robust and interpretable framework for automated fake news detection. Recommendations for platform developers, policymakers and media literacy educators are provided.
Statistical Assessment of Maternal Health Outcomes in Nigeria
Admin
About This Research Topic Maternal health, defined as health during pregnancy, childbirth and postpartum period, remains central to global public health. Sustainable Development Goal 3.1 targets global maternal mortality ratio below 70 per 100,000 live births by 2030, yet sub-Saharan Africa accounts for 66% of global maternal deaths. Nigeria, with less than 3% of world population, accounts for approximately 20% of global maternal deaths, with maternal mortality ratio of 512 per 100,000 live births in 2018 Nigeria Demographic and Health Survey, up from 576 in 2013 but far from required 7.5% annual reduction. Beyond mortality, maternal near-miss — woman who nearly died but survived life-threatening complication within 42 days of pregnancy termination — represents broader continuum. WHO near-miss criteria include cardiovascular, respiratory, renal, coagulation, neurological and uterine dysfunction, operationalisable from hospital records. Near-miss provides larger sample than death alone, enabling more powerful inference about risk factors. In Nigeria, direct causes include postpartum haemorrhage, hypertensive disorders, sepsis, obstructed labour and unsafe abortion, with indirect contributors malaria, anaemia and HIV. This study presents comprehensive biostatistical assessment using secondary data from 2018 NDHS and 1,200 maternal case records from University College Hospital, Ibadan spanning 2015-2023. Methods include descriptive statistics, chi-square and correlation, binary logistic regression for adverse outcome (near-miss or death), Kaplan-Meier and Cox proportional hazards for time-to-complication, ANOVA for birth weight across parity, and Principal Component Analysis for dimensionality reduction. UCH Ibadan serves large referral catchment across Oyo, Ogun and Osun, offering window into south-western Nigeria when combined with nationally representative NDHS. For foundational methods, see our guides to logistic regression and survival analysis in health research. Main Abstract Maternal mortality and morbidity remain pressing challenges in Nigeria, accounting for ~20% of global maternal deaths. This study presents statistical assessment using secondary data from 2018 Nigeria Demographic and Health Survey and 1,200 maternal case records from University College Hospital, Ibadan 2015-2023. Descriptive statistics characterized obstetric and sociodemographic distributions. Correlation and chi-square identified associations. Binary logistic regression modeled probability of adverse maternal outcome (maternal near-miss or death), survival analysis via Kaplan-Meier estimator and Cox proportional hazards assessed time-to-complication, ANOVA compared mean birth weights across parity groups, and PCA reduced dimensionality among correlated risk indicators. Findings: maternal age >35 years OR 2.84 (95% CI 1.97-4.10), absence of antenatal care OR 4.21 (2.93-6.05), referral delivery OR 3.17 (2.12-4.74), grand multiparity OR 2.51 (1.74-3.62), postpartum haemorrhage OR 6.83 (4.51-10.34) strongest independent predictors of adverse outcome. Cox model identified same variables as significant hazard contributors. Kaplan-Meier curves showed significant divergence in complication-free survival between women with and without ANC (log-rank p<0.001). PCA revealed two principal components explaining 61.3% variance. ANOVA confirmed significant difference in mean birth weight across parity groups (F=12.47, p<0.001). Findings support targeted ANC scale-up, skilled birth attendance improvement and emergency obstetric care strengthening.
Statistical Analysis of Mental Health Among University Students
Admin
About This Research Topic Mental health among university students has moved from peripheral concern to central public health priority. Globally, young adults aged 18 to 24 carry disproportionate burden of depression and anxiety, and university environments amplify vulnerability through academic pressure, financial strain, identity transitions and loss of familiar support. In Nigeria, underfunding, overcrowded facilities, calendar disruptions and escalating costs intensify these stressors, while counselling services remain thinly resourced and stigma suppresses help-seeking. This study provides comprehensive statistical analysis of mental health status among 350 undergraduates at a Nigerian federal university using the Depression, Anxiety and Stress Scale-21 (DASS-21). DASS-21 offers validated, brief measurement of three interrelated dimensions, suitable for non-clinical student populations and cross-cultural use. Methods include descriptive epidemiology, chi-square, independent t-tests, one-way ANOVA, Pearson and Spearman correlations, and multiple linear regression to move beyond prevalence counts to identification of high-risk subgroups and modifiable predictors. Findings show 47.4% of students with moderate to extremely severe symptoms on at least one subscale, with anxiety highest at 48.9%. Financial difficulty, academic workload perception and social support inadequacy emerge as strongest predictors, explaining 54.6% of variance. The article demonstrates a reproducible pipeline from descriptive to multivariate analysis, useful for student affairs planning and policy advocacy. For methodological guidance, see our resources on questionnaire design and regression analysis for social science research. Main Abstract Mental health disorders among university students have reached alarming levels globally, with depression, anxiety and stress affecting academic performance and wellbeing. This study undertakes comprehensive statistical analysis among 350 undergraduates at a Nigerian federal university using DASS-21 plus structured questionnaire capturing demographic, academic, financial and social variables. Methods included descriptive statistics, chi-square tests, independent t-tests, one-way ANOVA, Pearson and Spearman correlations, and multiple linear regression. Prevalence: 47.4% showed moderate to extremely severe symptoms on at least one DASS-21 subscale; depression 41.7%, anxiety 48.9%, stress 38.3%. Female students recorded significantly higher mean anxiety than males (t=-3.82, p<0.001). ANOVA revealed significant differences in mean stress across academic levels (F=4.63, p=0.003). Multiple regression identified financial difficulty (beta=0.341, p<0.001), academic workload perception (beta=0.287, p<0.001) and social support inadequacy (beta=-0.219, p=0.002) as strongest predictors of overall mental health score, model explaining 54.6% variance (Adjusted R2=0.546). Study concludes mental health distress is widespread and significantly predicted by structural, financial and social factors amenable to institutional intervention. Recommendations target university management, counselling centres and education policymakers.
Statistical Analysis of Malaria Incidence Among Rural Households
Admin
About This Research Topic Malaria continues to exact its heaviest toll in rural sub-Saharan Africa, where preventive interventions, housing quality and health literacy intersect to shape household risk. In Nigeria, which accounts for 27% of global cases, rural children under five experience malaria prevalence more than twice that of urban peers. Understanding why some households experience repeated episodes while others remain relatively protected requires household-level statistical analysis that goes beyond facility aggregates. This study examines malaria incidence among 320 rural households in Lere and Kachia Local Government Areas of Kaduna State between January 2022 and December 2023. It applies descriptive statistics, chi-square association tests, Pearson correlation, and count-data regression — Poisson and negative binomial — plus logistic regression for severe malaria. By integrating household survey data with primary health care records, the analysis quantifies incidence at 2.87 episodes per household per year and isolates modifiable predictors such as insecticide-treated net use, proximity to stagnant water, window screening and presence of children under five. The work demonstrates how appropriate count-data methods address overdispersion common in epidemiological counts, providing a methodological template for similar endemic settings. For students learning to model disease counts, our guides to Poisson regression and public health data analysis explain when Poisson assumptions fail and why negative binomial often fits better. Main Abstract Malaria remains leading cause of morbidity in rural Nigeria, yet household-level statistical analyses remain scarce. This study conducted comprehensive analysis of malaria incidence among rural households in Lere and Kachia LGAs, Kaduna State, using 320 households surveyed January 2022 to December 2023. Cross-sectional design with structured 28-item household questionnaire triangulated with PHC records. Descriptive statistics, chi-square tests, Pearson correlation, Poisson regression, negative binomial regression and logistic regression were applied. Overall incidence was 2.87 episodes per household per year (95% CI: 2.61-3.13). Households using insecticide-treated nets had significantly lower mean incidence than non-users (1.94 vs 3.81 episodes, p < 0.001). Poisson regression identified proximity to stagnant water (IRR=1.84, p<0.001), absence of ITN use (IRR=1.71, p<0.001) and presence of children under five (IRR=1.53, p<0.001) as significant predictors. Negative binomial model fitted overdispersed count data better (AIC=1,842.3 vs 2,104.7 Poisson). Logistic regression identified same factors plus lack of window screens as predictors of severe malaria. Findings support intensified ITN distribution, stagnant water drainage campaigns and community-based surveillance in rural Kaduna.
Explainable AI and Statistical Interpretation of Machine Learning Models
Admin
About This Research Topic Machine learning now influences decisions that affect health, credit, and justice, yet the most accurate models are often the least transparent. Clinicians, regulators and citizens increasingly ask not just how well a model predicts, but why it predicts that way. This question sits at the core of Explainable Artificial Intelligence (XAI), a field that seeks to make black-box models auditable, trustworthy and scientifically useful. This article presents a statistically grounded investigation of XAI using a real healthcare classification problem. Using 1,000 patient records from the UCI Heart Disease Repository, we train four classifiers — Logistic Regression, Random Forest, Gradient Boosting and Support Vector Machine — and interrogate them with SHAP, LIME and Partial Dependence Plots. Rather than treating explainability as a purely visual exercise, we apply hypothesis testing, correlation analysis and distributional checks to evaluate whether different explanation methods agree and whether they align with classical multivariate regression. The approach demonstrates how traditional statistical rigour can validate modern machine learning interpretability, a perspective especially relevant for statistics students learning to work with machine learning workflows. For readers building foundational skills, understanding exploratory data analysis, logistic regression and model evaluation metrics provides essential context before layering XAI techniques, which we cover in our guides to statistical modelling and machine learning project methods. Main Abstract This study examines the statistical interpretation of machine learning models through Explainable Artificial Intelligence (XAI). The opacity of high-performing algorithms limits responsible adoption in high-stakes domains where transparency, fairness and auditability are required. Using 1,000 records from the UCI Heart Disease (Cleveland) dataset, we conduct descriptive statistics, correlation and multivariate exploratory analysis, then train Logistic Regression, Random Forest, Gradient Boosting and Support Vector Machine classifiers. XAI tools — SHapley Additive exPlanations (SHAP), Local Interpretable Model-agnostic Explanations (LIME) and Partial Dependence Plots (PDP) — are applied to explain model behaviour. Statistical hypothesis tests compare predictive accuracy (AUC) across models and assess consistency between SHAP-derived feature importance rankings and logistic regression coefficient magnitudes. Random Forest achieved the highest accuracy at 91.4%, with significant differences in AUC across models (p < 0.05). SHAP importance rankings correlated significantly with classical regression coefficients, and age, serum cholesterol, maximum heart rate achieved, and resting blood pressure consistently emerged as top predictors across methods. PDPs confirmed clinically plausible marginal effects without evidence of spurious artefactual relationships. Findings affirm that XAI bridges statistical theory and machine learning practice, enabling validation of model decisions against domain knowledge. The study recommends mandatory integration of XAI into machine learning pipelines deployed in healthcare and other regulated sectors, supported by formal statistical validation.
Statistical Analysis of Climate Change Effects on Crop Yield in Benue State
Elijah T
About This Research Topic Benue State feeds Nigeria. Producing roughly 60% of the nation's yams and significant shares of rice, maize and sorghum, it employs over 80% of its workforce in agriculture. Yet the very climate that makes the Guinea Savannah productive is shifting. Farmers across its 23 local government areas now report later onset of rains, shorter growing seasons, and more frequent floods like those in 2012 and 2022. This study provides a 39-year statistical assessment of those shifts, using secondary data from 1985 to 2023 on annual rainfall and mean temperature from the Nigerian Meteorological Agency Makurdi Station and crop yields from the Benue State Agricultural Development Programme. Rather than relying on anecdote, it applies a complete inferential chain: descriptive decade analysis, Mann-Kendall trend testing with Sen's slope, Pearson and Spearman correlation, simple and multiple linear regression, and classical additive time series decomposition. The approach matters because most existing Nigerian studies use national aggregates or single-crop correlations without trend significance testing. By combining non-parametric trend detection robust to outliers with regression models that isolate joint effects of rainfall and temperature, this analysis delivers evidence that extension services and policymakers can act on. Readers unfamiliar with climate-agriculture linkages can start with our guide to climate change adaptation strategies in African agriculture. Main Abstract This research conducts a comprehensive statistical evaluation of climate change impacts on major crop yields in Benue State, Nigeria, known as the Food Basket of the Nation. Utilizing 39 years of annual observations from 1985 to 2023, the study integrates climate data on total annual rainfall and mean annual temperature from NiMet Makurdi and yield data for maize, rice, sorghum and yam from BNARDA. Descriptive analysis shows mean annual rainfall declined from 1,487 mm in 1985-1994 to 1,312 mm in 2015-2023, a drop of 175 mm (11.8%), while mean temperature rose from 27.1°C to 28.4°C (+1.3°C). Non-parametric Mann-Kendall tests confirm a significant declining rainfall trend (Kendall's tau = -0.341, p = 0.008, Sen's slope = -4.87 mm/year) and a significant rising temperature trend (tau = 0.487, p < 0.001, slope = 0.034°C/year). Pearson correlation indicates significant positive associations between rainfall and yields of maize (r = 0.712), rice (r = 0.689), sorghum (r = 0.624) and yam (r = 0.541), all p < 0.001, while temperature correlates negatively with all four crops. Multiple linear regression with rainfall and temperature as joint predictors explains 63.7% of maize yield variance (F = 31.49, p < 0.001), 57.8% for rice, 48.7% for sorghum and 41.2% for yam, with both predictors independently significant. Time series decomposition reveals declining trend components for maize, rice and sorghum, accelerating post-2010. The findings support promotion of drought-tolerant varieties, smallholder irrigation expansion, and strengthened agrometeorological advisory systems in Benue State.
Telemedicine Adoption and Healthcare Outcomes in Abuja
Elijah T
About This Research Topic Abuja has more going for telemedicine than almost anywhere else in Nigeria: strong internet coverage, a large educated middle class, and the headquarters of the very agencies shaping the country's digital health policy. Yet even here, adoption doesn't look the same everywhere. Step outside the well-connected core of Abuja Municipal Area Council into the peri-urban communities of Bwari, and telemedicine use drops noticeably, a gap that says a lot about who benefits from digital health in Nigeria today and who gets left behind. This article draws on a statistical study of telemedicine adoption and healthcare outcomes across AMAC and Bwari Area Council, tracing the journey from simple awareness through to regular use, and testing whether that use actually changes healthcare behaviour, not just attitudes. For readers curious how a study like this is structured statistically, our sample research projects library includes comparable quantitative studies worth reviewing as models. The findings matter for more than just Abuja. They speak to a question the whole country is grappling with: as telemedicine platforms multiply across Nigeria, who is actually adopting them, what's holding others back, and does regular use translate into measurable healthcare benefits or just convenience for people who were already going to be fine? The sections below walk through the background, the study's approach, and what the results suggest for policy and platform design going forward. Main Abstract Telemedicine, healthcare delivered through digital communication technologies, moved from a niche service to global prominence during the COVID-19 pandemic and now stands as one of the more significant shifts in how healthcare gets delivered. Nigeria's telemedicine ecosystem has grown steadily since 2015, with the pandemic sharply accelerating adoption. Even so, what actually drives people to adopt telemedicine, how far that adoption has spread beyond early, tech-savvy users, and whether regular use produces measurable healthcare benefits have remained poorly understood in Nigeria's urban context. This study statistically analysed telemedicine adoption patterns and healthcare outcomes among 374 adults across Abuja Municipal Area Council (AMAC) and Bwari Area Council in the Federal Capital Territory. It used a cross-sectional survey design alongside an independent samples t-test comparing healthcare utilisation between telemedicine users and non-users, framed by the Technology Acceptance Model and Diffusion of Innovations Theory. Telemedicine awareness reached 71.7 percent across the combined sample, with 48.1 percent having tried it at least once and 34.8 percent counting as regular users, AMAC consistently outperforming Bwari at every stage of adoption. Among regular users, satisfaction was high, with 78.5 percent reporting satisfaction and an average score of 3.92 out of 5. Most regular users, 71.5 percent, said telemedicine had genuinely improved how they managed a chronic condition, and 88.5 percent said they would recommend it to others. An independent samples t-test found that regular telemedicine users made significantly fewer unnecessary outpatient visits per year than non-users, 1.84 versus 3.41 visits on average, a large and clinically meaningful reduction. Chi-square tests confirmed that both smartphone ownership and internet quality were significantly associated with regular use, and logistic regression identified smartphone ownership, internet quality, prior positive digital health experience, and age as significant independent predictors of regular telemedicine adoption, while gender showed no significant effect. Based on these findings, the study recommends accelerating 4G and 5G broadband deployment in peri-urban FCT communities, integrating telemedicine into the NHIA benefit package, designing senior-citizen-friendly telemedicine interfaces, and running a national telemedicine literacy campaign. Together, these recommendations offer a concrete, evidence-based path toward expanding telemedicine adoption and capturing its health system efficiency benefits more broadly across Nigeria.
Health Insurance Utilization in Rivers State, Nigeria
Elijah T
About This Research Topic Rivers State is one of Nigeria's wealthiest states, powered by oil and gas revenue and a thriving urban economy in Port Harcourt. Yet wealth alone hasn't solved its health insurance problem. Most households in the state, like most households across Nigeria, still pay for healthcare out of their own pockets, exposed to exactly the kind of financial shock that health insurance is meant to prevent. That contradiction, high income sitting alongside low coverage, is what this article sets out to explain. Drawing on a household survey conducted across Port Harcourt Municipal and Obio-Akpor Local Government Areas, this piece looks at who actually enrols in health insurance, who doesn't, why, and, just as importantly, whether the people who do enrol actually use their coverage. For readers interested in how this kind of household-level economic research is put together, our sample research projects library includes comparable studies across economics and related disciplines. The findings carry weight well beyond Rivers State. They speak directly to Nigeria's push toward Universal Health Coverage and to a policy question that keeps resurfacing nationally: if income isn't the main barrier, what is? The sections below unpack the background to the problem, the study's approach, and what it suggests for closing the coverage gap. Main Abstract Health insurance is one of the central mechanisms for achieving Universal Health Coverage, protecting households from the kind of catastrophic financial hardship that a sudden illness can bring. Despite the creation of Nigeria's National Health Insurance Authority and various state-level schemes, enrolment nationally remains strikingly low, with fewer than 5 percent of Nigerians covered by any formal health insurance. Rivers State, despite ranking among Nigeria's wealthiest oil-producing states, mirrors this national picture: coverage clusters almost entirely within formal employment, leaving the much larger informal sector almost entirely uninsured. This study examined the socioeconomic, attitudinal, and structural factors shaping health insurance enrolment and utilization among 348 households in Port Harcourt Municipal and Obio-Akpor Local Government Areas of Rivers State. Using a cross-sectional survey design, a structured 32-item household questionnaire was administered between October and December 2024. The study looked at two distinct outcomes: whether a household had enrolled in any formal health insurance scheme, and whether enrolled households had actually used their covered services in the past year. Of the households surveyed, 120 (34.5 percent) had some form of health insurance enrolment. Among the 228 uninsured households, the inability to afford premium contributions was the single biggest barrier (43.0 percent), followed by simply not knowing which schemes were available (31.6 percent). Even among the 120 enrolled households, only 73 (60.8 percent) had actually used their covered services in the past 12 months, held back mainly by long facility waiting times (41.3 percent) and poor drug availability (30.4 percent). Logistic regression identified formal sector employment, household monthly income, awareness of NHIA schemes, perceived quality of care, and education level as significant independent predictors of enrolment, while gender showed no significant effect once other factors were accounted for. A willingness-to-pay analysis found that households were, on average, willing to pay less per month than the current formal sector premium charged by the Rivers State Contributory Health Commission, pointing to a subsidy gap that a targeted government subsidy could realistically close. Based on these findings, the study recommends mandatory employer-based enrolment across formal sector businesses, a subsidised enrolment pathway for informal sector workers, quality improvements at NHIA-accredited facilities, and a targeted awareness campaign reaching informal workers in markets, motor parks, and religious institutions. Together, these recommendations offer a data-driven route toward expanding health insurance coverage in Rivers State and moving closer to the Universal Health Coverage goal.
Childhood Immunization Coverage in Kano State, Nigeria
Elijah T
About This Research Topic Vaccines rank among the most effective tools medicine has ever produced, capable of preventing millions of childhood deaths every year. Yet in Kano State, Nigeria's most populous state, a large share of children still miss out on the full course of routine vaccinations, and the gap between rural and urban households remains wide. Understanding exactly where and why that gap persists is the difference between a vaccination programme that guesses and one that targets its resources where they matter most. This article draws on a statistical study of childhood immunization coverage in Kano Municipal (urban) and Kura (rural) Local Government Areas, examining coverage rates for each vaccine, dropout patterns, and the household factors that most strongly predict whether a child gets fully immunized. For readers curious how research like this is built from the ground up, our sample research projects library includes comparable statistical and public health studies worth reviewing as models. The findings matter well beyond the two Local Government Areas studied. They speak to a much larger question facing Nigeria's north-west zone and, by extension, the country's chances of reaching its 90 percent immunization coverage target. The sections that follow set out the background to the problem, the study's approach, and what its results suggest for closing the immunization gap. Main Abstract Immunization ranks among history's most cost-effective public health interventions, with the potential to prevent 4 to 5 million deaths every year from diseases that vaccines can stop. Yet Nigeria remains one of the countries with the highest number of unimmunized children in the world, and within Nigeria, the north-west zone, including Kano State, records some of the lowest coverage rates nationally. This study set out to statistically analyse childhood immunization coverage across rural and urban communities in Kano State, pinpointing which vaccines are most under-administered, how wide the rural-urban gap runs, and which household-level factors best explain why some children complete their full immunization schedule and others do not. Using a cross-sectional, community-based design and three-stage cluster sampling, the study collected data from 388 children aged 12 to 23 months and their mothers or caregivers, drawn from Kano Municipal LGA (urban) and Kura LGA (rural) between January and March 2024. Immunization status was confirmed through vaccination cards where available, supplemented by maternal recall. Coverage was calculated for all eight antigens in the WHO Expanded Programme on Immunisation schedule, and logistic regression was applied to identify which factors independently predicted full immunization. Overall full immunization coverage came to 48.9 percent, with a striking 22.5 percentage point gap between urban children (61.7 percent) and rural children (39.1 percent). BCG had the highest coverage of any single antigen at 91.7 percent, while Meningitis C trailed at 54.2 percent. The DPT/Penta series dropout rate reached 23.3 percent, more than double the WHO's recommended ceiling of 10 percent. Chi-square testing confirmed statistically significant links between full immunization and maternal education, distance to the nearest primary health care facility, household wealth, and LGA of residence. Logistic regression identified maternal knowledge of the vaccination schedule as the single strongest predictor of full immunization, followed by maternal education, household wealth, and distance to the nearest health facility. Urban residence remained significant even after adjusting for these other factors, while maternal age dropped out as a meaningful predictor once other variables were controlled for. The resulting model performed strongly, correctly classifying immunization status in the large majority of cases. Based on these findings, the study recommends stepped-up mobile vaccination outreach for communities more than 5 kilometres from a health facility, community mobilisation through religious leaders, structured prenatal education on the vaccination schedule delivered by community health workers, solar-powered cold chain equipment for rural facilities, and active follow-up of children who miss scheduled doses. Together, these measures target the specific determinants the study identified and offer a realistic path toward Kano State's 90 percent coverage goal.
Predicting Hospital Readmission Rates with Machine Learning
Elijah T
About This Research Topic Nearly one in five patients discharged from a Nigerian teaching hospital in this study ended up back on a ward within 30 days. That number is not unusual by global standards, but what is unusual is that almost none of Nigeria's tertiary hospitals currently use any data-driven tool to flag which patients are most likely to be readmitted before they walk out the door. This article works through a study that built and compared five statistical and machine learning models on patient discharge data from the University of Nigeria Teaching Hospital, Enugu, to see which approach best predicts 30-day unplanned readmission. Readers researching related quantitative or methodological topics can browse the Statistics project collection on ScholarNest for comparable studies in applied statistics and predictive modelling. What follows covers the background to hospital readmission research, the problem this study addresses, its objectives and key terms, and closes with frequently asked questions for students working on healthcare analytics. Main Abstract Hospital readmissions within 30 days of discharge represent one of the most widely used indicators of healthcare quality and patient safety, and impose substantial financial, clinical, and emotional burdens on patients, families, and healthcare systems. In Nigeria, despite the significant patient safety and resource implications of preventable readmissions, systematic data-driven approaches to identifying at-risk patients before discharge remain virtually absent from clinical practice. This study developed and compared five statistical and machine learning models for predicting 30-day unplanned hospital readmission using de-identified patient discharge data from the University of Nigeria Teaching Hospital (UNTH), Ituku-Ozalla, Enugu State, covering January 2020 to December 2023. A retrospective cohort design was adopted. De-identified discharge records of 1,200 adult inpatient admissions were extracted, covering cardiovascular, endocrine/metabolic, respiratory, renal, haematological, and other medical conditions. Key predictor variables included age, sex, primary diagnosis category, length of initial hospital stay, Charlson Comorbidity Index score, number of inpatient admissions in the 12 months preceding the index admission, discharge disposition, and whether a follow-up outpatient appointment was scheduled at discharge. The outcome variable was 30-day unplanned readmission, coded as a binary variable. The overall 30-day readmission rate in the sample was 18.7% (n = 224 readmissions out of 1,200 admissions, 95% CI: 16.5% to 20.9%). Cardiovascular disease patients had the highest readmission rate (25.8%), followed by endocrine and metabolic conditions including diabetes mellitus (22.4%). A strong dose-response relationship was identified between number of prior admissions and readmission rate, rising from 8.6% for patients with no prior admissions to 42.3% for those with three or more prior admissions in the past year. Five predictive models were developed and evaluated on a 30% holdout test dataset (n = 360): Binary Logistic Regression, Lasso-penalised Logistic Regression, Decision Tree, Random Forest (200 trees), and Gradient Boosting Machine (XGBoost implementation). The Gradient Boosting Machine achieved the highest performance with accuracy 86.7%, AUC-ROC 0.847, and F1-score 0.742, followed by Random Forest (AUC = 0.831, F1 = 0.718). Standard logistic regression achieved AUC = 0.782. Differences in model AUC-ROC were statistically significant as confirmed by DeLong's test (p < 0.05 for GBM versus logistic regression), leading to rejection of the corresponding null hypothesis. Logistic regression identified number of prior admissions (OR = 1.628, p < 0.001), against-medical-advice discharge (OR = 2.438, p < 0.001), absence of scheduled follow-up appointment (OR = 1.844, p < 0.001), and Charlson Comorbidity Index score (OR = 1.273, p < 0.001) as the four strongest independent predictors of readmission. Gender was not a significant predictor (p = 0.427). Gradient Boosting Machine feature importance analysis confirmed prior admissions (28.4%), Charlson Comorbidity Index score (19.9%), and absence of follow-up appointment (16.2%) as collectively accounting for 64.5% of model predictive power. The study recommends implementing Gradient Boosting Machine-based readmission risk scoring at discharge, mandating follow-up appointment scheduling for all high-risk patients, establishing post-discharge telephone follow-up programmes, and developing a national readmission reduction initiative within Nigeria's healthcare quality improvement framework.
Machine Learning Algorithms for Credit Risk Prediction
Elijah T
About This Research Topic Ask a lender which model best predicts loan default and the honest answer is: it depends on what you are optimising for. This study put seven classification algorithms head-to-head on the same dataset of loan applicants and found that the algorithm with the best overall accuracy was not the one with the best ability to catch actual defaulters, a distinction that matters enormously once a model moves from a spreadsheet into a real lending decision. This article works through that comparison, testing Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, K-Nearest Neighbours, Naive Bayes, and Support Vector Machine against 1,000 loan applicant records under realistic class-imbalanced conditions. Readers researching related quantitative or financial topics can browse the Business Administration project collection on ScholarNest for comparable studies in finance and risk management. What follows covers the background to credit risk modelling, the specific problem this study addresses, its objectives, questions, and hypotheses, the key statistical terms used throughout, and closes with frequently asked questions for students and researchers working on credit scoring and classification. Main Abstract Credit risk assessment remains a foundational function of financial institutions, directly determining loan approval decisions, pricing, and provisioning for expected credit losses, with the accuracy of default prediction models carrying direct implications for institutional profitability, financial stability, and, at a systemic level, the soundness of the broader financial system. This study statistically examined the determinants of loan default and comparatively evaluated seven machine learning classification algorithms, Logistic Regression, Decision Tree, Random Forest, Gradient Boosting, K-Nearest Neighbours, Naive Bayes, and Support Vector Machine, using a dataset of 1,000 loan applicant records comprising demographic, financial, and credit-history variables. The study determined the overall default rate and described applicant characteristics, examined bivariate relationships between candidate risk factors and default status, identified statistically significant predictors of default using binary logistic regression, and comparatively evaluated the classification performance of the seven algorithms under class-imbalanced conditions typical of credit risk data. It applied descriptive statistics, Pearson correlation, independent samples t-tests, one-way ANOVA, Chi-square tests, binary logistic regression, and a comprehensive seven-algorithm classifier comparison using accuracy, precision, recall, F1-score, and area under the ROC curve, with class-balanced weighting applied during model training and a stratified 70:30 train-test split used for classifier evaluation. Results showed an overall default rate of 28.50% (285 of 1,000 applicants). Previous default history (χ2 = 21.16, p < 0.001), collateral provision (χ2 = 4.40, p = 0.036), and employment status (χ2 = 15.25, p < 0.001) were all significantly associated with default. Defaulting applicants had significantly lower mean credit scores (576.36) than non-defaulting applicants (640.15; t = -12.03, p < 0.001) and significantly higher debt-to-income ratios (0.349 versus 0.217; t = 12.92, p < 0.001). Binary logistic regression identified credit score (OR = 0.304 per standard deviation, p < 0.001), debt-to-income ratio (OR = 3.601 per standard deviation, p < 0.001), employment years (OR = 0.614, p < 0.001), previous default history (OR = 2.887, p < 0.001), absence of collateral (OR = 1.859, p = 0.001), unemployment status (OR = 3.416, p < 0.001), and loan amount (OR = 1.397, p < 0.001) as statistically significant predictors of default, with the overall model explaining a substantial proportion of variance (McFadden pseudo R² = 0.342). In the seven-algorithm comparative classification exercise, Random Forest achieved the highest overall accuracy (81.67%) and precision (69.14%), while Logistic Regression achieved the highest area under the ROC curve (87.63%) and tied for the highest recall (79.07%) alongside Support Vector Machine, with Gradient Boosting, Naive Bayes, Decision Tree, and K-Nearest Neighbours trailing across most metrics. Random Forest and Gradient Boosting feature importance rankings both converged on debt-to-income ratio and credit score as the two dominant predictors, corroborating the logistic regression findings. The study concludes that debt-to-income ratio, credit score, prior default history, and employment stability are the dominant statistical determinants of credit default risk, and that the choice between Logistic Regression and Random Forest as a deployed credit-scoring model should be guided by an institution's relative prioritisation of overall classification accuracy, favouring Random Forest, against default-detection sensitivity and regulatory interpretability, favouring Logistic Regression, rather than by accuracy alone. It recommends that financial institutions incorporate debt-to-income ratio and credit score as primary automated screening criteria, apply class-rebalancing techniques when training credit-scoring models, and maintain logistic regression-based scorecards alongside ensemble methods to satisfy both predictive performance and regulatory interpretability requirements.
Predicting Student Academic Performance with Machine Learning
Elijah T
About This Research Topic More algorithms does not automatically mean better predictions. That is the uncomfortable finding at the centre of this study: a simple, interpretable logistic regression model outperformed five more complex machine learning algorithms, including random forest and support vector machines, at the job of flagging students likely to underperform. This article works through a study that compared multiple linear regression against six classification algorithms on a dataset of 800 student records, testing which approach actually identifies at-risk students most reliably. Readers exploring related quantitative research can browse the Statistics project collection on ScholarNest for comparable studies in applied statistics and data analysis. What follows covers the background to educational data mining, the specific problem this study addresses, its objectives, questions, and hypotheses, the key statistical terms used throughout, and closes with frequently asked questions for students and researchers working on predictive modelling in education. Main Abstract Predicting student academic performance using statistical and machine learning techniques has become an increasingly important area of educational data mining, offering institutions a way to identify at-risk students early and design targeted interventions to improve learning outcomes. This study statistically examined the determinants of student academic performance and comparatively evaluated multiple linear regression against five machine learning classification algorithms, Decision Tree, Random Forest, K-Nearest Neighbours, Naive Bayes, and Support Vector Machine, alongside binary logistic regression, using a dataset of 800 student records comprising demographic, socioeconomic, behavioural, and academic-history variables. The study described the sample's demographic and academic characteristics, examined bivariate relationships between candidate predictors and academic performance, identified statistically significant predictors of continuous final examination scores through multiple linear regression, and compared the classification performance of six algorithms in predicting high versus low academic performance. It applied descriptive statistics, Pearson correlation, independent samples t-tests, one-way ANOVA, multiple linear regression, binary logistic regression, and a comprehensive machine learning classifier comparison using accuracy, precision, recall, F1-score, and area under the ROC curve, with data partitioned into a stratified 70:30 train-test split for the classification exercise. Results showed a mean final examination score of 55.38 (SD = 12.88, out of 100). Multiple linear regression identified previous GPA (B = 10.709, p < 0.001), attendance rate (B = 0.376, p < 0.001), weekly study hours (B = 0.611, p < 0.001), tutoring support (B = 2.939, p < 0.001), internet access (B = 1.787, p = 0.010), sleep hours (B = 0.482, p = 0.043), and middle socioeconomic status relative to low (B = 1.919, p = 0.007) as statistically significant predictors of final examination score, with the overall model explaining 56.2 percent of score variance (R² = 0.562, Adjusted R² = 0.555, F = 77.69, p < 0.001). In the classification exercise predicting high versus low academic performance using a median split, logistic regression achieved the highest overall performance among all six algorithms evaluated (accuracy = 80.83%, AUC = 86.65%), outperforming the more algorithmically complex Random Forest (accuracy = 77.50%, AUC = 84.53%), Naive Bayes (accuracy = 77.08%, AUC = 83.98%), Support Vector Machine (accuracy = 76.25%, AUC = 82.72%), Decision Tree (accuracy = 72.50%, AUC = 75.70%), and K-Nearest Neighbours (accuracy = 65.00%, AUC = 72.08%). Random forest variable importance and decision tree feature importance both consistently identified previous GPA and attendance rate as the two dominant predictors, corroborating the multiple regression findings. The study concludes that prior academic achievement and class attendance are the most robust statistical determinants of subsequent academic performance, that logistic regression's strong comparative performance demonstrates that algorithmic complexity does not guarantee superior predictive accuracy, particularly where predictor-outcome relationships are approximately linear and additive, and that combining interpretable regression-based models with ensemble machine learning methods offers complementary value for educational early-warning system design. It recommends that educational institutions prioritise attendance monitoring and early academic support, particularly for students with weak prior academic records, and adopt interpretable models such as logistic regression as a first-line analytical tool before investing in more computationally intensive machine learning infrastructure.
STATISTICAL ANALYSIS OF FACTORS INFLUENCING VACCINE ACCEPTANCE AMONG ADULTS IN IBADAN METROPOLIS, NIGERIA
Admin
About This Research Topic Why do some adults roll up their sleeves for a vaccine without hesitation, while others delay, deliberate, or refuse outright, even when the vaccine is free and readily available? This article reworks a full undergraduate research project — titled Statistical Analysis of Factors Influencing Vaccine Acceptance Among Adults in Ibadan Metropolis, Nigeria — into a clear, search-friendly guide that keeps the original study's aim, objectives, and scope fully intact while making the material easier to read and more useful to search. Ibadan, as Nigeria's largest city by land area, offers a revealing lens on a question that matters well beyond its boundaries: what actually moves the needle, so to speak, on adult vaccine uptake in a large, diverse Nigerian urban centre. If you are shaping a related public health or biostatistics project of your own, it may help to browse similar statistics and public health project topics before finalising your title. The sections below walk through the background, problem statement, objectives, research questions, significance, scope, and key terms of the original study, followed by a set of frequently asked questions distilled from its findings on trust, safety perceptions, education, and social influence. Main Abstract Vaccine hesitancy ranks among the most consequential threats to public health worldwide, eroding immunisation coverage and creating the conditions for vaccine-preventable diseases to resurface. This study carried out a quantitative statistical investigation into the factors shaping vaccine acceptance among adults living in Ibadan Metropolis, Oyo State, Nigeria. Using a cross-sectional design, the researchers drew a sample of 300 adult residents through stratified random sampling across three Local Government Areas and administered a validated 28-item structured questionnaire covering vaccine acceptance, sociodemographic background, health beliefs, trust in healthcare providers, social influence, perceptions of vaccine safety, and sources of health information. The collected data were examined using descriptive statistics, chi-square tests, binary logistic regression, one-way ANOVA, and Pearson correlation. Findings showed that 67.0% of respondents accepted vaccines outright, 21.3% expressed hesitancy, and 11.7% refused vaccination altogether. The binary logistic regression model identified trust in healthcare providers (OR = 4.21, 95% CI: 2.67–6.64, p < 0.001), perceived vaccine safety (OR = 3.17, 95% CI: 2.01–5.00, p < 0.001), educational attainment (OR = 2.84, 95% CI: 1.74–4.64, p < 0.001), and social influence (OR = 2.31, 95% CI: 1.48–3.61, p < 0.001) as the strongest independent predictors of acceptance. One-way ANOVA detected significant differences in acceptance scores across educational levels (F = 14.37, p < 0.001) and income groups (F = 9.82, p < 0.001), while chi-square analysis confirmed significant associations between vaccine acceptance and both religious affiliation (chi-square = 18.43, df = 3, p < 0.001) and prior adverse vaccine experience (chi-square = 22.17, df = 1, p < 0.001). The study concludes that vaccine acceptance in Ibadan is a multidimensional outcome, shaped jointly by trust, perceived safety, education, social norms, and access to reliable information. It recommends targeted community engagement, stronger capacity building for healthcare workers, and culturally sensitive communication campaigns as practical levers for improving vaccine uptake in Ibadan and comparable Nigerian urban settings.
Deep Learning vs Classical Statistical Models: A Guide to Prediction Accuracy
Admin
About This Research Topic For students, researchers, and working analysts, the decision to build a predictive model around deep learning or around a classical statistical technique is no longer a purely academic exercise. It shapes how accurately a hospital can flag a suspicious diagnosis, how fairly a lender can assess a loan applicant, and how efficiently a school can identify a struggling student before it is too late. This article reworks a full undergraduate research project — titled Deep Learning versus Classical Statistical Models in Prediction Accuracy — into an accessible, search-friendly guide that keeps every element of the original research intact: its aim, its objectives, its research questions, and its scope. If you are scoping out a similar comparative modelling project of your own, you may find it useful to browse related statistics and data science project topics before settling on a final title. The central question this project investigates is deceptively simple: does deep learning actually outperform classical statistical models such as logistic regression, linear discriminant analysis, and naive Bayes when the task is predicting an outcome from data? As you will see in the sections below, the honest answer is 'it depends' — and the value of a rigorous comparative study lies precisely in mapping out what it depends on. Main Abstract Predictive analytics increasingly sits at the intersection of two competing traditions: the decades-old discipline of classical statistical modelling and the rapidly expanding field of deep learning. As artificial intelligence tools spread into healthcare diagnostics, credit assessment, education, and public policy, the question of which family of models delivers the most trustworthy predictions has taken on real practical weight. This project undertakes a disciplined, side-by-side evaluation of three classical statistical models — logistic regression, linear discriminant analysis, and naive Bayes — against three deep learning architectures — the multilayer perceptron, a convolutional neural network adapted for tabular data, and a long short-term memory network. The comparison draws on secondary data from three widely used benchmark sources: the Wisconsin Breast Cancer Diagnostic Dataset, the UCI Adult Income Dataset, and a credit risk dataset from the financial domain. Each model's performance is measured across a broad set of indicators — overall accuracy, sensitivity, specificity, area under the ROC curve, F1 score, root mean squared error, mean absolute error, and calibration error — and the differences between model families are tested formally using paired t-tests, the Wilcoxon signed-rank test, analysis of variance, and the Friedman non-parametric test. The evidence shows that deep learning architectures, particularly the multilayer perceptron, pull ahead of classical models on datasets that are large, high-dimensional, or shaped by non-linear relationships among predictors. On smaller, cleaner datasets, however, logistic regression holds its own and in some cases matches deep learning performance outright. A further, often overlooked finding is that deep learning models tend to be poorly calibrated relative to classical alternatives and demand considerably more computing power to train. The project concludes that model choice should follow from the specific characteristics of the dataset at hand rather than from prevailing trends, and it proposes a structured framework to guide that choice in practice.
Statistical Evaluation of AI-Based Disease Diagnosis Systems
Admin
About This Research Topic Modern healthcare is undergoing a profound shift as artificial intelligence takes on a growing share of diagnostic decision-making once reserved almost exclusively for clinicians. From reading medical images to flagging early markers of diabetes and cardiovascular disease, AI-powered diagnostic tools promise faster, more consistent, and more scalable disease detection. Yet impressive headline accuracy figures can conceal a more complicated reality: how a model performs on average is not the same as how it performs for every patient, every subgroup, or every clinical setting. This article unpacks a rigorous statistical evaluation of AI-based disease diagnosis systems, built around two widely used benchmark datasets. Rather than relying on a single accuracy score, the evaluation draws on a battery of statistical tools: descriptive statistics, logistic regression baselines, ROC and AUC analysis, calibration diagnostics, and subgroup fairness testing using the McNemar test. The result is a clearer picture of what AI diagnostic systems actually deliver — and where they fall short. Students and researchers working on similar quantitative or health-analytics topics may also find it useful to browse related statistics and data science project topics for inspiration on structuring their own evaluation studies. Main Abstract The rapid adoption of artificial intelligence in clinical diagnosis has outpaced the statistical scrutiny needed to confirm that these systems are safe, fair, and clinically dependable. This study carries out a wide-ranging statistical assessment of AI-based disease diagnosis systems, focusing on classification accuracy, sensitivity, specificity, the diagnostic odds ratio, area under the receiver operating characteristic curve (AUC-ROC), and calibration performance. Using secondary data drawn from two well-established public health datasets — the Pima Indians Diabetes Dataset and the UCI Heart Disease Dataset — the analysis combines descriptive statistics, hypothesis testing, correlation analysis, logistic regression, ROC curve analysis, and the McNemar test to compare model performance. Convolutional neural networks and gradient boosting classifiers outperformed a conventional logistic regression baseline, recording AUC-ROC scores of 0.91 and 0.89 respectively, against 0.76 for logistic regression. Despite this aggregate advantage, the study uncovered statistically significant disparities in sensitivity and specificity across demographic subgroups, alongside evidence of mild overconfidence in the models' probability estimates. These findings indicate that although AI diagnostic tools can outperform traditional statistical baselines on average, subgroup performance gaps and calibration weaknesses must be addressed before wide-scale clinical rollout. The study recommends fairness-aware training methods, ongoing statistical performance audits, and the incorporation of explainability techniques into clinical validation workflows.
Can't find your topic? Request a custom material →
