Project Marketplace

Project Topics & Materials

Search curated materials across every major department, faculty and institution.

Showing 20 of 25 materials

ZERO-TRUST ARCHITECTURE IMPLEMENTATION MODEL FOR SMALL BUSINESS NETWORKSComputer Science

ZERO-TRUST ARCHITECTURE IMPLEMENTATION MODEL FOR SMALL BUSINESS NETWORKS

Scholarnesthub Admin

About This Research Topic For much of history of enterprise networking information security architecture organised around trusted internal network perimeter defended by firewalls, VPNs, intrusion detection systems positioned at boundary on implicit assumption any user or device successfully authenticated onto internal network can subsequently be trusted to interact relatively freely with internal resources. This perimeter-based trust but verify once model reasonably well matched to era employees worked from fixed office locations using company-owned devices connecting to on-premises servers. Zero-Trust Architecture implementation model for small business networks Operating environment has changed substantially. Proliferation of cloud-hosted business applications, normalisation of remote and hybrid work arrangements, and widespread adoption of BYOD practices mean traditional network perimeter has become porous, distributed, and in many organisations difficult to meaningfully define. Small businesses particularly affected: constrained IT budgets and limited dedicated security staffing mean many small business networks continue to rely on legacy perimeter-based controls single firewall and VPN even as actual usage patterns cloud application access, remote work, personal device use have moved well beyond assumptions model depends. Zero-Trust Architecture responds to mismatch by discarding assumption of implicit trust based on network location entirely. Under Zero-Trust model formalised authoritatively in NIST Special Publication 800-207, every access request regardless of whether originates inside or outside traditional perimeter evaluated per-request against policy considering requesting identity, device posture, sensitivity of requested resource, and relevant contextual signals with access granted on least-privilege basis for specific transaction rather than standing grant of broad network access. While conceptual case for Zero-Trust well established and increasingly mandated in large enterprise and government contexts, practical implementation historically associated with level of architectural complexity and licensing cost placing comprehensive ZTA adoption largely out of reach for small business IT environments operating with limited budgets and generalist rather than security-specialist technical staff. This study addresses gap by designing, implementing, and evaluating Zero-Trust implementation model specifically scoped and cost-engineered for small business network constraints built predominantly on open-source and low-cost commercial components and organised around phased adoption roadmap intended to make meaningful Zero-Trust security gains achievable without budget and staffing levels typically associated with enterprise ZTA deployments. Main Abstract Small businesses increasingly rely on cloud-hosted applications, remote and hybrid work arrangements, and bring-your-own-device (BYOD) practices, yet typically continue to rely on legacy perimeter-based network security models that implicitly trust any user or device once granted access to the internal network. This structural mismatch leaves small business networks disproportionately exposed to credential-based intrusion, lateral movement, and ransomware propagation, while the substantial cost and complexity historically associated with Zero-Trust security models have placed such architectures largely out of reach for resource-constrained small business information technology environments. This study presents the design, implementation, and evaluation of a practical, cost-conscious Zero-Trust Architecture (ZTA) implementation model tailored specifically to small business network constraints, grounded in the National Institute of Standards and Technology (NIST) Special Publication 800-207 Zero Trust Architecture reference framework. The study adopted a Design Science Research methodology to translate the seven NIST ZTA tenets into a concrete, implementable reference architecture comprising identity-centric access control, device posture verification, micro-segmentation, and continuous policy evaluation, implemented using predominantly open-source and low-cost commercial components suited to small business budgets, including a self-hosted identity provider, a software-defined perimeter/policy engine, and endpoint posture agents. A functional testbed was implemented, simulating a representative 25-endpoint small business network spanning cloud application access, an on-premises file server, and remote/BYOD endpoints, and was evaluated against an equivalent conventional perimeter/VPN-based baseline network in terms of lateral movement containment, unauthorised access prevention, policy enforcement latency, and administrative overhead. Results showed that the Zero-Trust testbed successfully contained 47 of 50 simulated lateral movement attempts following an initial single-endpoint compromise (94% containment), compared to 11 of 50 (22% containment) for the conventional baseline network, while introducing a modest average authentication/authorisation latency overhead of 180 milliseconds per access request, assessed as acceptable relative to the security gain. Usability and administrability evaluation conducted with 15 respondents (small business IT administrators and staff users) using a Likert-scale questionnaire yielded a mean System Usability Scale (SUS)-equivalent score of 72.4. The study concludes that a right-sized, phased Zero-Trust implementation model, built substantially on open-source tooling, is both technically achievable and operationally viable for small business networks, and provides a phased adoption roadmap intended to lower the practical barrier to Zero-Trust adoption in resource-constrained organisational contexts. Keywords: Zero-Trust Architecture, NIST SP 800-207, micro-segmentation, identity-centric security, small business networks, lateral movement containment, software-defined perimeter

₦5,000View
Predictive Analytics Model for Student Dropout Risk in Tertiary InstitutionsComputer Science

Predictive Analytics Model for Student Dropout Risk in Tertiary Institutions

Scholarnesthub Admin

About This Research Topic Student attrition — withdrawal from tertiary academic programme prior to completion whether through voluntary withdrawal academic exclusion or extended ultimately unresumed leave — represents substantial persistent challenge for higher education globally. Beyond personal and economic cost borne by affected student whose interrupted education may leave them with debt or foregone earnings without corresponding qualification attrition represents direct institutional cost: resources invested in recruitment enrolment early support not recouped through eventual graduation and elevated attrition rates can affect funding accreditation standing reputation. At SCHOLARNESTHUB, we transform educational data mining research into SEO-optimized academic resources. This study on predictive analytics model for student dropout risk is crafted for students searching for education project topics and computer science project topics . Institutional responses historically predominantly reactive: advising and support interventions frequently triggered only after student exhibited overt signs of academic distress failing grades extended absence formal withdrawal application by which point range of effective intervention narrowed relative to what might have been possible with earlier identification. Educational data mining and learning analytics research over past decade increasingly explored whether routinely collected institutional data — demographic information captured at admission socioeconomic indicators such as tuition payment status and scholarship/loan status and early academic performance records — can be used to construct predictive models capable of identifying individual students at elevated dropout risk substantially earlier in programme than conventional reactive advising typically allows creating window for proactive rather than reactive intervention. ML offers natural approach: given historical dataset of past students demographic socioeconomic academic records together with eventual outcome graduated still enrolled or dropped out classifier can learn statistical patterns associated with eventual dropout and subsequently applied to current students to generate individual risk estimate. Practical value depends critically on ability to generate sufficiently accurate predictions using only information available early enough for meaningful intervention — constraint this study addresses directly by restricting input features to information available by end of second semester. Main Abstract Student attrition remains persistent and costly challenge for tertiary institutions worldwide representing both loss of individual educational and economic opportunity for affected student and loss of institutional resources invested in enrolment and early academic support. Early identification of students at elevated risk would in principle allow institutional academic support services to intervene proactively — through advising financial aid counselling or academic remediation — rather than reactively after student already withdrawn or academically excluded. This study presents design implementation and evaluation of predictive analytics model for estimating individual student dropout risk using historical academic demographic and socioeconomic data to identify students at elevated risk early enough in academic programme for meaningful institutional intervention to remain possible. Study adopted Design Science Research methodology structuring development around data preprocessing, feature engineering, model training, and evaluation stages. Four classification algorithms implemented and comparatively evaluated: Logistic Regression (as interpretable baseline), Random Forest, Gradient Boosting (XGBoost), and fully connected Deep Neural Network trained on combined dataset comprising publicly available UCI/Kaggle “Predict Students' Dropout and Academic Success” dataset (4,424 records from Portuguese higher education institution spanning demographic socioeconomic and first- and second-semester academic performance features) supplemented with synthetically constructed Nigerian-context dataset of 2,000 records constructed to reflect institution-reported aggregate dropout statistics and enrolment demographic distributions from three Nigerian universities given unavailability of comparable individual-level public Nigerian datasets. Models evaluated on ability to predict dropout status using only information available by end of student's second semester reflecting practical constraint that any useful early-warning system must generate predictions early enough for intervention to remain meaningful. Results showed XGBoost achieved strongest predictive performance (accuracy = 87.9%, F1-score = 0.869, AUC = 0.932) followed by Random Forest (86.4% accuracy) and DNN (85.1% accuracy) with Logistic Regression trailing (81.2% accuracy) but providing directly interpretable coefficient-based risk factors. Feature importance analysis identified first- and second-semester grade performance, number of enrolled versus approved curricular units, age at enrolment, and tuition payment status as most predictive features with socioeconomic and financial features collectively contributing substantial share of predictive power alongside purely academic performance indicators. Supplementary threshold-sensitivity analysis examined precision-recall trade-off at varying risk-classification thresholds relevant to institutional decisions about intervention resource allocation. Study concludes machine learning-based dropout risk prediction using data already routinely collected by tertiary institutions can identify at-risk students with sufficient accuracy and lead time to support proactive institutional intervention and recommends institutional pilot deployment as early-warning decision-support tool integrated with existing academic advising workflows with appropriate attention to fairness and non-punitive use of risk predictions.

₦5,000View
Intrusion Detection System Using Machine Learning for Network Traffic AnalysisComputer Science

Intrusion Detection System Using Machine Learning for Network Traffic Analysis

Scholarnesthub Admin

About This Research Topic Network Intrusion Detection Systems (NIDS) are foundational component of organisational cybersecurity infrastructure tasked with monitoring network traffic to identify patterns indicative of malicious activity whether originating from external attacker attempting to breach perimeter or compromised internal host engaged in lateral movement data exfiltration or participation in broader distributed attack. NIDS complement endpoint-focused controls (antivirus, EDR) by providing visibility into network-layer activity not observable from single endpoint vantage and detecting attack classes — network scanning reconnaissance denial-of-service flooding distributed botnet coordination — whose malicious character most readily apparent from aggregate traffic pattern analysis rather than single-host monitoring alone. At SCHOLARNESTHUB, we transform cybersecurity research into SEO-optimized academic resources. This study on intrusion detection system using machine learning for network traffic analysis is crafted for students searching for computer science project topics and cybersecurity project topics . Conventional signature-based NIDS operate by matching observed traffic against database of known attack signatures specific byte patterns packet sequences or protocol anomalies previously catalogued. While computationally efficient and highly precise against previously catalogued patterns this shares structural limitation well documented for signature-based malware detection: inability to detect novel techniques and persistent detection gap during interval between new technique emergence and signature cataloguing and distribution. Machine learning-based intrusion detection offers complementary paradigm learning to distinguish benign from malicious based on statistical and behavioural features extracted from network flow data — aggregated characteristics of sequence of packets sharing common source destination protocol such as flow duration packet count byte count inter-arrival time statistics — rather than matching exact byte-level signatures. Model trained on sufficiently representative diverse labelled dataset can in principle generalise to detect attack traffic exhibiting statistical characteristics similar to training distribution even where specific tool or exact packet sequence differs. This study designs implements and comparatively evaluates multiple ML algorithms for flow-based NIDS addressing both binary and multi-class tasks using two widely adopted public benchmarks with goal of providing empirically grounded guidance on which approach offers most favourable accuracy class-balance robustness and inference-latency trade-off for practical deployment. Main Abstract Network intrusion detection remains foundational component of organisational cybersecurity infrastructure tasked with identifying malicious or anomalous network traffic indicative of ongoing attack whether originating from external adversary or compromised internal host. Conventional signature-based Network Intrusion Detection Systems (NIDS) which match observed traffic patterns against database of known attack signatures share same structural limitation documented extensively for signature-based malware detection: inability to detect novel attack patterns not yet catalogued in signature database. Machine learning-based intrusion detection which learns to distinguish benign from malicious traffic based on statistical and behavioural features extracted from network flow data offers complementary potentially more generalisable detection paradigm. This study presents design implementation and evaluation of machine learning-based NIDS operating on flow-level network traffic features comparing multiple classification algorithms to identify most effective approach for both binary (benign/malicious) and multi-class (specific attack category) classification tasks. Study adopted Design Science Research methodology structuring development around feature extraction, model training, and comparative evaluation stages. Four classification algorithms implemented and comparatively evaluated: Random Forest, Support Vector Machine (SVM), fully connected Deep Neural Network (DNN), and ensemble Gradient Boosting model (XGBoost), each trained and evaluated on widely used CICIDS2017 and NSL-KDD benchmark intrusion detection datasets comprising flow-level statistical features (packet counts, byte counts, flow duration, inter-arrival time statistics, and TCP flag distributions) extracted from labelled network traffic captures spanning both benign traffic and multiple attack categories including Denial-of-Service (DoS), Distributed Denial-of-Service (DDoS), port scanning/reconnaissance, brute-force credential attacks, and web application attacks. Results showed XGBoost achieved highest overall performance for binary classification (accuracy = 99.1%, F1-score = 0.990, AUC = 0.998), followed closely by Random Forest (98.6% accuracy) and DNN (98.2% accuracy), with SVM trailing (95.7% accuracy) but remaining viable lower-complexity alternative. For multi-class attack-category classification XGBoost again achieved strongest performance (96.3% macro-averaged accuracy), though all models showed comparatively reduced recall specifically for minority attack classes (notably web application attacks and subset of rare DoS variant subtypes) reflecting pronounced class imbalance characteristic of source datasets. Feature importance analysis identified flow duration, total forward packets, and destination port as consistently among most discriminative features across models. Comparative inference latency benchmarking found XGBoost capable of classifying network flows at throughput suitable for near-real-time deployment on standard network monitoring hardware. Study concludes ensemble tree-based methods and gradient boosting specifically offer most favourable accuracy-latency trade-off for flow-based ML NIDS among architectures evaluated and recommends targeted class-imbalance mitigation techniques to improve minority-class attack detection in future deployment-oriented work.

₦5,000View
MALWARE CLASSIFICATION USING DEEP LEARNING ON EXECUTABLE FILE FEATURESComputer Science

MALWARE CLASSIFICATION USING DEEP LEARNING ON EXECUTABLE FILE FEATURES

Scholarnesthub Admin

About This Research Topic Malicious software encompassing viruses, worms, trojans, ransomware, spyware designed to disrupt, damage, or gain unauthorised access has grown substantially over past two decades driven partly by commoditisation of malware-development toolkits and partly by increasing use of obfuscation, polymorphism automatically varying byte-level signature, and packing compressing or encrypting payload to conceal content. Malware classification using deep learning on executable file features Conventional antivirus detection historically relied predominantly on signature-based matching comparing file byte content or hash against maintained database of known malicious signatures. While efficient and accurate against previously catalogued threats, signature-based detection structurally reactive: only detects variants already observed, analysed, added to database, leaving gap against novel zero-day malware and variants deliberately engineered through polymorphism or minor modification to evade exact matching while retaining same malicious functionality. Machine learning-based malware classification offers complementary paradigm: rather than matching exact signatures, model trained to recognise structural and behavioural patterns statistically associated with malicious intent learned from labelled training corpus, with goal of generalising to previously unseen variants sharing underlying characteristics even where exact byte-level signature differs. Deep learning approaches given capacity to automatically learn hierarchical feature representations from comparatively raw input rather than requiring extensive manual feature engineering have shown particular promise recent research, applied both to structured features extracted from executable's file format metadata and in some approaches directly to raw byte-level file content. This study focuses specifically on static analysis-based deep learning classification — features extracted from executable without executing it — in contrast to dynamic analysis executing sample within instrumented sandbox to observe runtime behaviour. Static offers operational advantage avoiding infrastructure cost, execution risk, latency associated with sandboxed dynamic execution, at potential cost of reduced effectiveness against malware employing runtime-only obfuscation not observable from static structure alone; trade-off examined empirically in evaluation. Main Abstract The continued proliferation and rapid evolution of malicious software (malware) presents a persistent challenge to conventional signature-based antivirus detection, which relies on matching a suspicious file against a database of known malicious byte-sequence signatures and is therefore structurally limited in its ability to detect novel, obfuscated, or polymorphic malware variants not yet represented in the signature database. Machine learning-based malware classification, which learns to distinguish malicious from benign executables based on extracted structural and behavioural features rather than exact signature matching, has emerged as a complementary detection approach with the potential to generalise to previously unseen malware variants. This study presents the design, implementation, and evaluation of a deep learning-based malware classification system operating on static features extracted from Windows Portable Executable (PE) files, without requiring dynamic execution of potentially malicious code. The study adopted a Design Science Research methodology, structuring development around feature extraction, model architecture design, training, and evaluation stages. Static features were extracted from each executable's PE header structure, section table, imported API function list, and byte-level entropy profile, yielding a structured feature vector per sample without executing any file. Three model architectures were implemented and comparatively evaluated: a baseline Random Forest classifier (as a non-deep-learning comparison baseline), a fully connected Deep Neural Network (DNN) operating on the structured PE feature vector, and a one-dimensional Convolutional Neural Network (CNN) operating directly on raw byte-level file content represented as a fixed-length byte sequence. All models were trained and evaluated on a combined dataset drawn from the publicly available EMBER (Endgame Malware BEnchmark for Research) feature dataset and the SOREL-20M malware corpus metadata, comprising 40,000 labelled samples (20,000 malicious, 20,000 benign) after balancing and preprocessing, with a stratified 70/15/15 train/validation/test split. Results showed that the CNN model achieved the highest classification accuracy (97.8%), followed by the DNN (96.3%) and the Random Forest baseline (94.1%), with the CNN also achieving the highest F1-score (0.978) and area under the ROC curve (0.991). Feature importance analysis of the Random Forest baseline and integrated-gradients attribution analysis of the DNN both identified imported API function patterns and section-table entropy characteristics as the most discriminative feature categories, consistent with known malware obfuscation and packing behaviour patterns documented in the malware analysis literature. Comparative inference-time benchmarking found the CNN's per-sample classification latency (14ms average) suitable for near-real-time endpoint scanning use cases. The study concludes that deep learning-based static analysis of executable file features offers a computationally efficient, execution-free complement to signature-based detection, achieving strong classification accuracy while avoiding the operational risk and latency associated with dynamic (sandboxed execution) analysis approaches, and recommends further research into ensemble combination of static deep-learning classification with complementary dynamic analysis techniques. Keywords: malware classification, deep learning, convolutional neural network, static analysis, Portable Executable, EMBER dataset, SOREL-20M, AUC 0.991

₦5,000View
SECURITY VULNERABILITY ASSESSMENT OF MOBILE BANKING APPLICATIONSComputer Science

SECURITY VULNERABILITY ASSESSMENT OF MOBILE BANKING APPLICATIONS

Scholarnesthub Admin

About This Research Topic The past decade has witnessed substantial shift in financial service delivery from physical banking halls toward mobile-first channels driven by rising smartphone penetration expansion financial inclusion initiatives and operational cost advantages mobile channels offer financial institutions. In Nigeria Central Bank Nigeria cashless policy and financial inclusion strategy have further accelerated adoption mobile banking applications which now handle substantial share retail transaction volume including funds transfer bill payment airtime purchase and account management functions previously restricted to banking hall or ATM interaction. This expansion in mobile banking usage has been accompanied by corresponding expansion in attack surface available to malicious actors. Mobile banking application particularly attractive target because it typically handles highly sensitive assets authentication credentials transaction PINs account balances ability to initiate irreversible funds transfers while operating in inherently less controlled environment than bank internal server infrastructure: application executes on user-owned device unknown security posture transmits data over networks varying trustworthiness including public Wi-Fi and distributed through application stores where reverse engineering installed package technically straightforward for motivated attacker. Industry and academic security research bodies most notably Open Worldwide Application Security Project OWASP have documented recurring categories mobile application vulnerability including insecure data storage weak server-side controls insufficient transport layer protection and reliance on client-side security controls that can be bypassed on rooted or jailbroken device through structured references such as OWASP Mobile Top 10 and Mobile Application Security Verification Standard MASVS. Despite availability standards empirical assessments deployed banking applications particularly in developing-economy contexts continue report presence well-documented avoidable vulnerability classes suggesting persistent gap between availability security guidance and consistent application in practice. Recent assessments include security maturity assessment Indonesian Android mobile banking apps using MobSF and OWASP MASVS Level 2 MASVS-R APK files Google Play Store analysed using Static Application Security Testing MobSF evaluated against OWASP MASVS and security evaluation mobile banking applications Sudan utilizing Static Application Security Testing SAST via MobSF Quixxi evaluated against OWASP MASVS where APK files obtained Google Play Store analysed using Static Application Security Testing Mobile Security Framework MobSF and evaluated against OWASP MASVS Level 2. This study conducts structured standards-aligned security vulnerability assessment mobile banking application functionality using combination static and dynamic analysis techniques applied within controlled ethically sound test environment with goal quantifying prevalence common vulnerability classes and producing actionable prioritised remediation guidance. For related materials see ScholarNestHub cybersecurity collection . Main Abstract Rapid adoption mobile banking applications across Nigeria and other developing economies has expanded financial inclusion but has simultaneously widened attack surface available to malicious actors targeting sensitive financial data and funds. Study presents structured security vulnerability assessment mobile banking applications aimed at identifying classifying quantifying common security weaknesses present in Android-based banking applications and at proposing mitigation guidelines aligned with recognised industry standards. Study adopted hybrid research methodology combining structured static and dynamic Mobile Application Security Testing MAST process guided by OWASP Mobile Application Security Verification Standard MASVS and Mobile Security Testing Guide MSTG with Design Science Research approach for developing accompanying automated assessment toolkit. Rather than targeting live production banking applications without authorisation which would be unlawful and unethical study constructed representative test application replicating common mobile banking functionality login balance enquiry funds transfer PIN/biometric authentication and additionally assessed curated set publicly available deliberately-vulnerable banking-style training applications OWASP mobile testing benchmark applications providing controlled ethically sound assessment environment. Static analysis performed using MobSF Mobile Security Framework and manual manifest/code review; dynamic analysis performed using Drozer and Burp Suite in controlled proxy/emulator environment. Total 12 vulnerability categories drawn from OWASP Mobile Top 10 were assessed across sample applications including insecure data storage weak cryptography insecure communication insufficient certificate pinning improper platform usage. Results showed 8 of 12 vulnerability categories were present in at least one assessed sample applications with insecure local data storage unencrypted shared preferences and SQLite databases and absent or misconfigured SSL/TLS certificate pinning identified as most prevalent weaknesses appearing in 75% and 62.5% of assessed samples respectively. Weighted risk-scoring model incorporating exploitability and impact dimensions consistent with OWASP Risk Rating Methodology applied to prioritise findings and set mitigation recommendations including mandatory use Android Keystore-backed encryption certificate pinning root/jailbreak detection and secure coding checklists compiled into practical remediation guideline. Study concludes mobile banking applications even where developed by financially resourced institutions commonly exhibit avoidable security weaknesses that structured standards-aligned assessment methodology can systematically surface and recommends financial institutions incorporate MASVS-aligned testing into software development lifecycle rather than relying solely on post-deployment penetration testing.

₦5,000View
COST-OPTIMIZATION FRAMEWORK FOR MULTI-CLOUD DEPLOYMENT STRATEGIESComputer Science

COST-OPTIMIZATION FRAMEWORK FOR MULTI-CLOUD DEPLOYMENT STRATEGIES

Scholarnesthub Admin

About This Research Topic Cloud computing has become dominant infrastructure paradigm with major public providers Amazon Web Services Microsoft Azure Google Cloud Platform offering extensive catalogue of compute storage managed services each governed by provider-specific pricing structures. Increasing number of organisations for reasons including avoidance of vendor lock-in compliance with data residency sovereignty regulatory requirements necessitating hosting within particular jurisdictions best served by different providers regional footprints and resilience against single provider outage have adopted multi-cloud deployment strategies in which overall workload portfolio distributed across two or more public providers simultaneously distinct from single-provider or hybrid-cloud combining public with on-premises. While multi-cloud offers strategic benefits it introduces substantial cost-management complexity. Each provider maintains distinct pricing model encompassing purchasing options on-demand pay-as-you-go reserved capacity commitments discounted rates in exchange for commitment period and spot/preemptible instances steep discounts on spare capacity subject to reclamation with limited notice region-specific variation and discount programmes not directly comparable without normalisation. Workload placement decision that appears cost-optimal within single provider catalogue may not be optimal across full cross-provider landscape yet many organisations placement decisions driven by technical or organisational convenience rather than systematic quantified optimisation. This cost challenge motivated emergence of Cloud Financial Operations FinOps as discipline focused on financial accountability and systematic optimisation of variable consumption-based spending. This article for SCHOLARNESTHUB presents rewritten SEO-optimized study designing implementing evaluating cost-optimization framework combining workload characterisation cross-provider price modelling and mixed-integer linear programming optimisation recommending cost-minimising placement subject to redundancy latency residency constraints. For related cloud research see cloud computing project topics on SCHOLARNESTHUB . Main Abstract Organisations increasingly deploy workloads across multiple public cloud providers simultaneously strategy termed multi-cloud deployment motivated by considerations including vendor lock-in avoidance geographic/regulatory data residency requirements and resilience against single-provider outages yet strategy introduces substantial cost-management complexity relative to single-provider deployment since each provider maintains distinct pricing models discount structures and billing granularity and workload placement decisions that appear cost-optimal in isolation for single provider may not be cost-optimal when evaluated against full multi-provider pricing landscape. Study presents design implementation and evaluation of cost-optimization framework for multi-cloud deployment strategies combining workload characterisation cross-provider price modelling and optimisation algorithm that recommends cost-minimising workload placement and purchasing-option configuration on-demand reserved capacity or spot/preemptible instances across multiple cloud providers subject to defined performance and availability constraints. Study adopted Design Science Research methodology structuring development around cost-model construction optimisation algorithm design comparative evaluation stages. Framework ingests workload resource utilisation profiles CPU memory storage network transfer and applies mixed-integer linear programming optimisation model to recommend for each workload component cost-minimising combination of cloud provider instance/service type region and purchasing option subject to constraints including minimum redundancy workload components must be distributed across at least two providers for defined critical services maximum acceptable latency and data residency requirements. Framework evaluated using publicly available pricing data from three major cloud providers Amazon Web Services Microsoft Azure Google Cloud Platform applied to three representative workload case studies of differing characteristics steady-state web application batch data-processing workload with predictable off-peak execution windows and bursty unpredictable-demand workload constructed to reflect realistic organisational usage patterns. Results showed optimisation framework recommended multi-cloud placement achieved 31.4 percent cost reduction relative to naive single-provider on-demand baseline deployment for steady-state web application workload 47.2 percent reduction for batch processing workload through aggressive use of spot/preemptible capacity during its flexible execution window and 22.8 percent reduction for bursty workload while satisfying all defined redundancy and latency constraints in each case. Sensitivity analysis varying redundancy constraint strictness quantified cost premium associated with stronger multi-provider resilience guarantees finding 12.3 percent average cost premium for upgrading from single-provider to dual-provider redundancy for critical workload components. Study concludes systematic optimisation-driven multi-cloud workload placement can achieve substantial cost reduction relative to common single-provider or ad hoc multi-cloud allocation practice while providing organisations explicit quantified visibility into cost-resilience trade-off inherent in multi-provider redundancy decisions and recommends framework adoption as decision-support tool for organisational FinOps practice.

₦5,000View
DATA-DRIVEN FRAMEWORK FOR OPTIMIZING PUBLIC TRANSPORTATION ROUTESComputer Science

DATA-DRIVEN FRAMEWORK FOR OPTIMIZING PUBLIC TRANSPORTATION ROUTES

Scholarnesthub Admin

About This Research Topic Efficient public transport is the backbone of urban productivity, yet in many fast-growing Nigerian cities, route networks were never formally designed. In Lagos, danfo minibuses, shared taxis and okada services dominate daily trips, with routes that emerged organically through driver habit and incremental demand rather than systematic planning. While this organic growth shows adaptability, it creates structural inefficiencies: overlapping corridors, excessive transfers, long passenger wait times, and unbalanced vehicle loads. Explore recent transport planning research topics This article presents a rewritten and expanded version of an undergraduate project on a data-driven framework for optimizing public transportation routes, retaining the original aim and methodology while adding academic depth, SEO clarity, and practical interpretation for transit authorities and researchers. Main Abstract Public transportation route networks in rapidly urbanizing cities, particularly in paratransit-dominated systems such as Nigerian minibus, shared taxi, and motorcycle taxi services, have often evolved incrementally without systematic optimization against current ridership demand. This leads to measurable inefficiencies including long wait times, uneven loading, and poor coverage. This study designs, implements, and evaluates a data-driven framework for route optimization that integrates ridership demand analysis with graph-based network optimization. The research adopts a Design Science Research methodology, organized into demand analysis, road network graph construction, genetic algorithm-based route design, and comparative evaluation. A synthetic transportation network modeled on a mid-sized Nigerian city district was built from OpenStreetMap data, supplemented with validation using the publicly available Chicago Transit Authority ridership dataset. The optimization uses a genetic algorithm with a multi-objective fitness function weighting total passenger travel time, vehicle operating cost proxied by total route distance and fleet size, and demand coverage defined as the proportion of origin-destination pairs served within an acceptable transfer threshold. An agent-based passenger simulation compared the optimized network against the baseline. Results indicate an 18.3% reduction in average passenger travel time from 42.6 to 34.8 minutes, a 24.1% improvement in average vehicle load factor, and a 9.7% reduction in total fleet-distance operating cost, while demand coverage increased from 89.6% to 94.2% within the transfer threshold. Sensitivity analysis characterizes the trade-off frontier between travel time and operating cost as fitness weights vary. The study concludes that genetic algorithm-based optimization can produce substantially more efficient configurations than incrementally evolved networks and recommends a phased pilot beginning with highest-impact routes.

₦5,000View
Big Data Framework for Real-Time Flood/Disaster Risk AnalyticsComputer Science

Big Data Framework for Real-Time Flood/Disaster Risk Analytics

Scholarnesthub Admin

About This Research Topic Flooding is among most frequently occurring and economically damaging categories of natural disaster worldwide and represents particularly recurrent seasonal hazard across Nigeria where riverine flooding along Niger and Benue systems and tributaries combined with rapid often poorly planned urban expansion into flood-prone low-lying areas resulted in repeated large-scale events causing substantial loss of life displacement and economic damage including widely documented 2012 and 2022 Nigerian flood events both affecting multiple states and displacing well over million people collectively. At SCHOLARNESTHUB, we transform disaster analytics and big data projects into SEO-optimized academic resources. This study on big data framework for real-time flood/disaster risk analytics is crafted for students searching for computer science project topics and environmental management project topics . Effective flood early-warning depends fundamentally on timely integration and analysis of multiple heterogeneous sources: river gauge telemetry measuring water level, rainfall station and satellite-derived precipitation estimates, soil-moisture and antecedent-condition data influencing runoff vs infiltration, and topographic/hydrological terrain determining inundation extent. These individually high-volume and high-velocity continuously updating collectively present integration and near-real-time processing challenge exceeding practical capacity of conventional single-server database and batch architectures and naturally suited to distributed stream-oriented big data approaches capable of ingesting processing analysing multiple concurrent streams with low latency. Existing flood monitoring in Nigeria coordinated primarily through Nigeria Hydrological Services Agency (NIHSA) historically constrained by limited real-time station density fragmented data integration across agencies and correspondingly limited lead time and geographic precision in risk communication. Big data stream-processing architectures combined with ML risk modelling offer viable approach to addressing integration and latency constraints enabling ingestion and near-real-time joint analysis to generate more timely geographically granular assessments. Main Abstract Flooding remains one of most frequently occurring and economically damaging natural disaster categories globally and particularly recurrent seasonal hazard across Nigeria's riverine and low-lying urban areas where limited real-time monitoring infrastructure and fragmented data sources historically constrained lead time and geographic precision of flood early-warning capability. Effective flood risk analytics requires integration and near-real-time processing of multiple heterogeneous high-volume high-velocity data streams — river gauge and rainfall sensor telemetry, satellite-derived precipitation and soil-moisture estimates, and topographic/hydrological terrain data — data integration and processing challenge naturally suited to big data architectural approaches rather than conventional single-server database and batch-processing techniques. This study presents design implementation and evaluation of big data framework for real-time flood and disaster risk analytics integrating distributed stream-processing pipeline, hydrological risk-scoring model, and geographic visualisation dashboard evaluated using combination of publicly available historical hydrological and satellite precipitation datasets and simulated real-time sensor telemetry stream constructed to reflect realistic river gauge and rainfall station reporting patterns for representative Nigerian river basin case study area. Study adopted Design Science Research methodology structuring development around distributed data ingestion, stream processing, risk modelling, and visualisation stages. System architecture integrates Apache Kafka for distributed stream ingestion, Apache Spark Structured Streaming for near-real-time data processing and aggregation, hydrological risk-scoring model combining threshold-exceedance rule component with Random Forest-based flood-likelihood classifier trained on historical river-level and rainfall antecedent-condition data, and geospatial dashboard rendering current risk status by sub-catchment area. System evaluated along three dimensions: end-to-end processing latency under simulated concurrent sensor load, flood risk classification accuracy benchmarked against historical documented flood event dates for case study basin, and system throughput scalability under increasing simulated sensor node count. Results showed streaming pipeline sustained median end-to-end latency 2.8 seconds from simulated sensor reading ingestion to updated risk-score availability under simulated load 500 concurrent virtual sensor nodes comfortably within sub-5-minute latency target considered operationally meaningful for flood early-warning use cases. Random Forest flood-likelihood classifier evaluated retrospectively against 15 years historical documented flood event dates achieved recall 88.2% for correctly flagging documented flood event periods as elevated risk with false-positive rate 9.4% for non-flood periods. Throughput scalability testing found architecture maintained stable processing latency up to 1000 simulated concurrent sensor nodes before exhibiting beginning of queueing-related latency degradation indicating headroom substantially beyond case study basin's actual monitoring station density. Study concludes distributed stream-processing big data architecture combined with hybrid threshold-and-machine-learning risk-scoring approach provides technically viable scalable foundation for real-time flood risk analytics suited to Nigerian river basin monitoring contexts and recommends phased pilot integration with Nigerian Hydrological Services Agency monitoring infrastructure as next step toward operational deployment.

₦5,000View
BIOMETRIC AUTHENTICATION SYSTEM COMBINING FINGERPRINT AND FACIAL RECOGNITIONComputer Science

BIOMETRIC AUTHENTICATION SYSTEM COMBINING FINGERPRINT AND FACIAL RECOGNITION

Scholarnesthub Admin

About This Research Topic Biometric authentication — verification of identity through measurable physiological or behavioural characteristics — has become pervasive component of contemporary access control ranging from smartphone unlock to border control, banking authentication, and physical facility access. Among available traits, fingerprint and facial recognition are two most widely deployed owing to comparatively low cost of capture hardware, broad user familiarity, and mature recognition algorithms. Biometric authentication combining fingerprint and facial recognition Unimodal systems however which rely on single trait carry inherent limitations. Fingerprint recognition can be degraded by worn, dirty, or injured fingertips and by residue-based presentation attacks lifted latent prints reproduced in gelatin or silicone documented extensively. Facial recognition while convenient and contactless sensitive to lighting, pose, expression, ageing, and vulnerable to presentation attacks using printed photographs, digital screen replay, or increasingly sophisticated synthetic media. Because unimodal decision rests entirely on single evidentiary source any weakness directly translates into exploitable vulnerability or legitimate-user failure. Multimodal biometric systems address limitation by combining evidence from two or more independent traits on premise that attacker capable of spoofing one modality substantially less likely to simultaneously spoof second independent modality while combination of complementary evidence also tends to improve overall accuracy provided fusion strategy well designed. Fusion can occur at several levels — sensor level, feature level, score level, or decision level — each carrying different trade-offs between complexity and achievable accuracy gain. Score-level fusion combining normalised match confidence scores produced by independent unimodal matchers widely regarded as offering favourable balance of tractability and empirical performance and is approach adopted in this study. This study designs, implements, and evaluates multimodal system combining fingerprint and facial recognition via score-level fusion with goal of empirically quantifying accuracy and spoofing-resistance improvement achievable relative to either unimodal approach using combination of established public datasets and live volunteer data collection. Main Abstract Unimodal biometric authentication systems, relying on a single biometric trait such as fingerprint or facial recognition alone, remain vulnerable to spoofing attacks, environmental sensor noise, and trait-specific failure conditions (fingerprint wear or occlusion; facial recognition sensitivity to lighting, pose, and presentation attacks) that can compromise both security and usability. Multimodal biometric systems, which combine two or more independent biometric traits, offer a means of mitigating these individual weaknesses through complementary evidence fusion, at the cost of increased system complexity. This study presents the design, implementation, and evaluation of a multimodal biometric authentication system combining fingerprint and facial recognition, employing a score-level fusion strategy to combine the outputs of independent unimodal matchers into a single authentication decision. The study adopted the Design Science Research methodology, structuring development around feature extraction, unimodal matching, score normalisation, and fusion stages. Fingerprint recognition was implemented using minutiae-based feature extraction (ridge ending and bifurcation points) with a minutiae-matching algorithm, while facial recognition was implemented using a deep convolutional neural network (a FaceNet-style embedding model) generating 128-dimensional facial embeddings compared via Euclidean distance. Match scores from both modalities were normalised using min-max normalisation and combined using a weighted sum fusion rule, with fusion weights optimised on a validation subset. The system was implemented as a functional prototype comprising an enrolment module, a fingerprint capture and matching pipeline, a facial capture and matching pipeline, and a fusion decision engine, and was evaluated using a combined dataset drawn from the publicly available SOCOFing fingerprint dataset and the Labelled Faces in the Wild (LFW) facial dataset, paired synthetically to construct 300 simulated multimodal user identities, supplemented with a live-capture evaluation set of 40 volunteer participants. Results showed that the multimodal fusion system achieved an Equal Error Rate (EER) of 0.9%, compared to 2.8% for the standalone fingerprint matcher and 3.4% for the standalone facial matcher, and achieved a False Acceptance Rate (FAR) of 0.3% at a False Rejection Rate (FRR) of 1.1% at the selected operating threshold, outperforming both unimodal baselines across the receiver operating characteristic curve. Spoofing resistance testing using printed photograph and video-replay presentation attacks against the facial modality, and using gelatin-mould fingerprint replicas against the fingerprint modality, found that the fusion system correctly rejected 96% of combined spoofing attempts, compared to 71% for the standalone facial matcher and 78% for the standalone fingerprint matcher, since a successful attack against the fused system required simultaneously defeating both modalities. Usability evaluation with 40 participants yielded a mean authentication time of 3.2 seconds and a System Usability Scale (SUS)-equivalent score of 80.1. The study concludes that score-level fusion of fingerprint and facial recognition meaningfully improves both authentication accuracy and spoofing resistance relative to either unimodal approach, at an acceptable usability cost, and recommends multimodal biometric authentication for access control contexts where security requirements justify the additional implementation complexity. Keywords: multimodal biometrics, fingerprint recognition, facial recognition, score-level fusion, spoofing resistance, equal error rate, FaceNet, minutiae

₦5,000View
CLOUD-BASED E-LEARNING MANAGEMENT SYSTEM WITH ADAPTIVE CONTENT DELIVERYComputer Science

CLOUD-BASED E-LEARNING MANAGEMENT SYSTEM WITH ADAPTIVE CONTENT DELIVERY

Scholarnesthub Admin

About This Research Topic E-learning management systems have become central component tertiary education delivery providing digital infrastructure for content distribution assessment academic record-keeping. Majority widely deployed learning management platforms including those commonly used within Nigerian tertiary institutions predominantly implement fixed-sequence content delivery model: all enrolled learners within course progress through identical pre-determined sequence instructional modules regardless individual differences prior knowledge learning pace demonstrated performance formative assessment encountered along way. This uniform delivery model administratively simple widely familiar but well documented in educational technology learning sciences literature as pedagogically suboptimal relative adaptive approaches that dynamically adjust content sequencing pacing difficulty in response learner demonstrated mastery since learners entering course with differing prior preparation differing in-course learning trajectories are under fixed-sequence model given no individually differentiated instructional response to these differences. Adaptive learning technology most rigorously grounded in mastery-based instructional models informed by learner-modelling techniques such as Bayesian Knowledge Tracing addresses limitation by continuously estimating each learner probability having mastered each specific instructional concept based on observed formative assessment response pattern and using estimate to dynamically select subsequent content advancing learner who demonstrated mastery toward more advanced material while directing learner exhibiting persistent difficulty specific concept toward supplementary remedial content targeted at that specific gap rather than proceeding uniformly regardless demonstrated need. Separately e-learning platforms deployed within resource-constrained institutional contexts including many Nigerian tertiary institutions face distinct but related practical challenge infrastructure scalability under variable and at certain periods notably examination weeks when concurrent platform usage typically spikes substantially above baseline high concurrent user load. Conventional fixed-capacity single-server or modestly provisioned on-premises hosting deployments prone to severe performance degradation outright unavailability precisely during high-stakes high-load periods well-documented operational pain point institutional e-learning platforms. Cloud-native architectural approaches incorporating auto-scaling infrastructure that dynamically provisions additional compute capacity in response measured load offer technically mature solution to scalability challenge widely adopted commercial e-learning platforms but less consistently adopted within resource-constrained institutional deployments which may lack cloud architecture expertise budget model suited elastic usage-based cloud infrastructure spending. This study designs implements evaluates cloud-based e-learning management system incorporating Bayesian Knowledge Tracing-informed adaptive content delivery engine addressing both pedagogical limitation uniform content delivery and infrastructure scalability challenge within single integrated system evaluated through controlled comparative learning-outcome study and infrastructure load-testing. Recent implementations include microservices architecture Bayesian knowledge tracing KT Service Python FastAPI Content AI BERT embeddings and AI-powered adaptive learning platform using Bayesian Knowledge Tracing four-parameter probabilistic model mastery probability real time adaptive quiz generation where Bayesian Knowledge Tracing four-parameter probabilistic model updates student mastery probability after every quiz answer real time and quizzes dynamically generated from student weakest knowledge components difficulty scaling mastery level. For related project materials see ScholarNestHub computer science education collection . Main Abstract Conventional e-learning management systems typically deliver fixed uniform sequence instructional content to all enrolled learners regardless individual differences prior knowledge learning pace performance formative assessment one-size-fits-all delivery model well documented educational technology literature as suboptimal relative to instructional approaches that adapt content sequencing difficulty to individual learner demonstrated mastery. Simultaneously e-learning platforms deployed within resource-constrained institutional contexts including many Nigerian tertiary institutions face practical infrastructure challenges related to variable internet connectivity server scalability under concurrent examination-period load and operational cost maintaining dedicated on-premises hosting infrastructure. Study presents design implementation evaluation cloud-based e-learning management system incorporating adaptive content delivery engine addressing both pedagogical limitation uniform content delivery and infrastructure scalability challenge through cloud-native architecture. Study adopted Design Science Research methodology structuring development around adaptive engine design cloud architecture design and comparative evaluation stages. Adaptive content delivery engine implements mastery-based sequencing algorithm informed by Bayesian Knowledge Tracing which estimates learner probability having mastered each instructional concept from formative quiz response history and dynamically selects next content unit difficulty level accordingly branching learners who demonstrate mastery toward more advanced content while directing learners exhibiting difficulty toward supplementary remedial content addressing specific concept not yet mastered. System implemented using microservices-based cloud architecture deployed on Amazon Web Services comprising independently scalable content-delivery adaptive-engine and assessment microservices behind auto-scaling load balancer with content stored in cloud object storage service and served via content delivery network to address connectivity-variability concerns. System evaluated along two dimensions: learning outcome effectiveness comparing group 45 volunteer student participants using adaptive delivery pathway against comparison group 43 participants using equivalent fixed-sequence delivery pathway covering identical instructional content both assessed via identical pre-test/post-test instrument; and infrastructure scalability load-testing cloud architecture under simulated concurrent user load representative institutional examination-period usage spike. Results showed adaptive-delivery group achieved significantly greater mean post-test score improvement mean gain 24.6 percentage points than fixed-sequence comparison group mean gain 17.1 percentage points t(86)=3.42 p<0.001 while also completing equivalent instructional content in significantly less average time 38.2 minutes versus 47.9 minutes t(86)=4.18 p<0.001. Scalability load-testing found auto-scaling microservices architecture maintained 95th-percentile page response time below 1.2 seconds under simulated load 5000 concurrent users scaling from baseline 2 to 14 application server instances during load ramp compared to fixed-capacity single-server baseline deployment which exhibited response time degradation exceeding 8 seconds and increasing request failure rate beyond approximately 800 concurrent users. Study concludes combining mastery-based adaptive content delivery with cloud-native auto-scaling microservices architecture measurably improves both learning outcomes infrastructure resilience relative to conventional fixed-sequence fixed-capacity e-learning deployment and recommends institutional adoption adaptive delivery techniques alongside cloud infrastructure modernisation for Nigerian tertiary institution e-learning platforms.

₦5,000View
CUSTOMER CHURN PREDICTION MODEL FOR TELECOM/FINTECH COMPANIESComputer Science

CUSTOMER CHURN PREDICTION MODEL FOR TELECOM/FINTECH COMPANIES

Scholarnesthub Admin

About This Research Topic Customer churn discontinuation of active relationship whether explicit cancellation non-renewal or sustained inactivity represents persistent financially significant challenge for subscription-based recurring-revenue businesses including telecommunications operators and fintech providers. Marketing literature long establishes cost of acquiring replacement customer substantially exceeds cost of retaining existing one making timely identification of at-risk customers commercially significant. Telecom operators face intense price competition number portability promotional switching incentives historically high churn rates. Fintech encompassing digital banking payments lending faces analogous churn dynamic where disengagement declining transaction activity dormancy precedes formal closure equally significant but more gradual. Predictive churn modelling applies classification to historical account usage billing engagement data to estimate likelihood of churning within future period enabling move from reactive undifferentiated retention spend toward proactive targeted intervention focused on elevated risk and when combined with customer lifetime value further prioritised toward highest-value at-risk where limited budget yields greatest protected revenue per intervention cost. This article for SCHOLARNESTHUB presents rewritten SEO-optimized study designing implementing and comparatively evaluating Logistic Regression Random Forest XGBoost and Deep Neural Network on Telco Customer Churn dataset 7,043 records and banking dataset 10,000 records with CLV-weighted evaluation. For similar analytics frameworks see data science project topics on SCHOLARNESTHUB . Main Abstract Customer churn discontinuation of customer relationship with service provider represents significant recurring revenue threat for subscription-based telecommunications and fintech businesses where cost of acquiring replacement substantially exceeds retention cost. Predictive churn modelling uses historical account usage engagement data to estimate individual likelihood of churning within defined future period enabling proactive targeted retention intervention rather than reactive undifferentiated spend. Study presents design implementation and evaluation of churn prediction model applicable to telecommunications and fintech contexts comparatively evaluating multiple classification algorithms and further examining customer lifetime value implications of applying churn model to prioritise retention resource allocation. Study adopted Design Science Research methodology structuring development around data preprocessing feature engineering model training evaluation stages. Four algorithms implemented and evaluated: Logistic Regression as interpretable baseline Random Forest XGBoost and fully connected Deep Neural Network trained on widely used publicly available Telco Customer Churn dataset 7,043 telecommunications customer records and supplementary publicly available bank customer churn dataset 10,000 retail banking records each comprising tenure service usage product holding billing demographic features with binary churn outcome label. Given moderate class imbalance churn representing approximately 26.5 percent and 20.4 percent of telecom and banking datasets respectively class-weighted training applied across all models with SMOTE oversampling additionally evaluated. Results showed XGBoost achieved strongest predictive performance on both datasets telecom accuracy 80.8 percent F1 0.612 AUC 0.847 banking accuracy 86.9 percent F1 0.634 AUC 0.872 followed closely by Random Forest with Logistic Regression trailing but providing directly interpretable coefficients and DNN performing comparably to but not exceeding tree-based ensembles. Feature importance analysis identified contract type tenure monthly charges as most predictive in telecom dataset and number of products held account activity status age as most predictive in banking dataset indicating meaningfully different churn driver profiles across sectors despite shared general approach. Supplementary customer lifetime value-weighted evaluation prioritising retention outreach toward highest-value at-risk customers identified by model rather than treating all flagged uniformly found prioritisation strategy could capture disproportionately large share of at-risk revenue within constrained retention-outreach budget relative to undifferentiated outreach strategy. Study concludes gradient boosting-based churn prediction combined with CLV-weighted retention prioritisation offers practically effective revenue-relevant approach to churn management for telecom fintech businesses while cautioning sector-specific feature importance differences indicate models should not be assumed directly transferable across sectors without retraining on sector-appropriate data.

₦5,000View
CYBERSECURITY AWARENESS AND THREAT PERCEPTION AMONG NIGERIAN UNIVERSITY STUDENTSComputer Science

CYBERSECURITY AWARENESS AND THREAT PERCEPTION AMONG NIGERIAN UNIVERSITY STUDENTS

Scholarnesthub Admin

About This Research Topic Nigerian universities are undergoing rapid digital transformation. Learning management systems, institutional portals, digital libraries, and mobile financial services now mediate almost every aspect of student life. While this connectivity expands educational opportunity, it also exposes students to an evolving landscape of cyber threats. Recent studies on information security behaviour in higher education show that undergraduates are frequently targeted because they combine high digital activity with limited formal security training. This article re-examines the original undergraduate project on Cybersecurity Awareness and Threat Perception Among Nigerian University Students, presenting a rewritten, expanded, and SEO-optimized academic discussion that retains the original objectives while deepening its analytical value. We synthesize findings from a survey of 368 students across three Nigerian universities to provide a clear diagnostic baseline for institutional intervention. Understanding where awareness is strong, where it is uneven, and where institutional communication fails is critical for protecting both students and university networks. Main Abstract University students constitute a high-exposure, high-interest group for cybersecurity research. They are intensive users of email, social media, cloud collaboration tools, and institutional information systems, yet unless enrolled in Computer Science or allied programmes, they rarely receive structured information security education. This gap between digital exposure and security competence increases vulnerability to phishing, social engineering, password compromise, and malware. This study assessed cybersecurity awareness and threat perception among undergraduate students in Nigerian universities and examined how academic and demographic factors relate to awareness levels. Using a descriptive cross-sectional survey design, a structured 28-item questionnaire measured four constructs: general cybersecurity awareness, phishing and social engineering threat perception, password and account security practices, and institutional and policy awareness. The instrument used a five-point Likert scale and was administered to a stratified random sample of 400 undergraduates drawn from Computer Science and non-Computer Science departments across three universities. After data cleaning, 368 valid responses were retained (92% response rate). Internal consistency was good (Cronbach’s alpha = 0.84). Descriptive results indicated a moderate overall awareness mean of 3.42 out of 5 (SD = 0.61). Password and account security practice recorded the highest construct mean (3.71), while institutional and policy awareness recorded the lowest (2.89). Inferential analysis showed Computer Science students scored significantly higher than non-Computer Science students (t(366) = 5.87, p < 0.001). Academic level had a significant effect (F(3, 364) = 4.21, p = 0.006), with final-year students outperforming first-year students. Prior self-reported exposure to phishing or social engineering attempts was positively correlated with overall awareness (r = 0.31, p < 0.001). The findings confirm moderate but uneven awareness, with institutional policy communication emerging as a critical weakness. The study recommends mandatory, recurring cybersecurity awareness training embedded in general orientation rather than confined to technical curricula.

₦5,000View
PHISHING WEBSITE DETECTION USING URL AND CONTENT-BASED FEATURE ANALYSISComputer Science

PHISHING WEBSITE DETECTION USING URL AND CONTENT-BASED FEATURE ANALYSIS

Scholarnesthub Admin

About This Research Topic Phishing remains among the most prevalent and financially damaging cyber-attack categories, exploiting deceptive websites impersonating legitimate services to harvest credentials. According to Anti-Phishing Working Group (APWG) trend reports , phishing consistently ranks among top reported attack vectors, with attackers increasingly leveraging short-lived, rapidly rotating domains that evade reactive defenses. Traditional blocklist-based browser warnings check visited URLs against databases of known malicious sites. While widely deployed, they are inherently reactive, unable to protect against newly registered zero-day phishing sites not yet catalogued. Machine-learning-based detection addresses this by learning characteristic patterns in URL structure, domain properties, and page content to enable proactive classification at point of access. This article presents a complete pipeline combining URL-lexical, domain/host-based and content-based features with explicit ablation quantification, packaged as low-latency detection service. For additional cybersecurity research materials, see ScholarNestHub cybersecurity collection and recent studies on hybrid feature-based phishing detection . Main Abstract Phishing remains one of the most prevalent and financially damaging cyber-attack categories exploiting deceptive websites impersonating legitimate services to harvest credentials and financial information, with industry reports ranking it among most reported attack categories. Blocklist-based defenses flagging known malicious sites remain widely deployed but inherently reactive, unable to protect against newly registered sites not yet catalogued. This study designs, implements and evaluates machine-learning-based phishing website detection system combining URL-lexical, domain/host-based and page-content features enabling proactive classification at point of access. Study adopted Design Science Research methodology combined with CRISP-DM for data-driven components. Combined dataset of 11,430 labelled websites constructed from PhiUSIIL phishing URL dataset and Kaggle-sourced phishing websites dataset incorporating 30 engineered features spanning three categories: URL-lexical (URL length, IP address presence, URL-shortening services, suspicious character counts), domain/host-based (domain age, WHOIS registration length, DGA-like pattern), and content-based (login form presence, ratio of external to internal links, favicon origin, mismatch between visible link text and target). Data cleaned and used to train and compare four models: Logistic Regression, Random Forest, XGBoost, and Multi-Layer Perceptron with feature-category ablation experiments isolating incremental contribution of URL-only, content-only and combined sets. XGBoost trained on combined feature set achieved strongest performance with accuracy 97.8%, precision 97.2%, recall 97.6%, F1-score 97.4%, outperforming URL-only subset (94.1% accuracy) and content-only subset (93.6% accuracy), demonstrating complementary rather than redundant discriminative signal. Trained model packaged as lightweight browser-extension-style detection service exposed via Flask backend evaluating visited page URL and rendered content in real time displaying risk indicator, achieving average end-to-end classification latency 140 milliseconds comfortably within range required for non-intrusive browsing. Study concludes combining URL-lexical and content-based features within gradient-boosted model provides materially more robust phishing detection capability than either category alone, and recommends periodic retraining and integration with live blocklist feeds as complementary safeguards. Keywords: phishing detection, machine learning, URL analysis, content-based features, cybersecurity, XGBoost, ablation study

₦5,000View
Explainable AI Framework for Medical Diagnosis Decision SupportComputer Science

Explainable AI Framework for Medical Diagnosis Decision Support

Scholarnesthub Admin

About This Research Topic Artificial intelligence has moved from a research curiosity into a quiet, everyday presence inside clinics and hospitals, supporting decisions on everything from cardiovascular risk to diabetes screening. Yet the models that tend to predict best, gradient-boosted ensembles and deep neural networks in particular, are often the hardest to interpret. A clinician handed a risk score with no supporting rationale is effectively being asked to trust a black box with a patient's welfare, and that is a proposition regulators, professional bodies, and clinicians themselves are increasingly unwilling to accept without question. This tension between predictive accuracy and interpretability sits at the centre of explainable AI (XAI) research in healthcare, and it is the problem this article works through in depth. Drawing on a completed applied research project, the discussion below walks through the design, implementation, and evaluation of an XAI framework for medical diagnosis decision support. Rather than treating explanation as an afterthought bolted onto a finished model, the underlying study places three complementary explanation techniques, SHAP, LIME, and rule extraction, alongside high-performing diagnostic classifiers built for heart disease and diabetes prediction, and then tests those explanations directly with practising clinicians and final-year medical students. That last step is what distinguishes this work from a great deal of published research in the field. Many studies apply an explanation technique to a medical model and treat the mere technical presence of that explanation as evidence of transparency, without ever asking the people who will actually use it whether the explanation makes clinical sense. The sections that follow set out exactly how this study answered that question, what it found, and what it means for anyone building, evaluating, or procuring a clinical decision-support tool. Main Abstract Machine learning models now deliver strong predictive performance across many medical diagnosis tasks, but the models that perform best, particularly ensemble tree-based methods and deep learning architectures, are often opaque, giving little indication of the reasoning behind any single prediction. That opacity is a genuine barrier to clinical adoption. Clinicians carry professional and ethical responsibility for diagnostic decisions, and a growing body of regulatory guidance calls for some form of explanation to accompany automated decision support in healthcare. This study responds to that gap by designing, building, and evaluating an explainable AI framework that pairs high-performing diagnostic classifiers with several complementary post-hoc explanation techniques, and by testing the resulting explanations empirically with clinical volunteers rather than relying on model accuracy alone as a stand-in for clinical usefulness. The project followed Design Science Research (DSR) methodology alongside the Cross-Industry Standard Process for Data Mining (CRISP-DM) for its data-driven components, drawing on two established public clinical datasets, the UCI Heart Disease dataset and the Pima Indians Diabetes dataset, together covering 1,536 patient records after cleaning. Each condition was modelled separately given the differing feature schemas. Logistic Regression, Random Forest, and XGBoost were trained and compared for each condition, and the best-performing model for each was paired with SHAP for global and instance-level feature attribution, LIME for local surrogate-model explanations, and a rule-extraction technique that produced human-readable if-then rules for clearly separable cases. XGBoost delivered the strongest diagnostic performance across both conditions, reaching 88.9% accuracy and an 87.2% F1-score for heart disease prediction, and 84.6% accuracy with a 78.3% F1-score for diabetes prediction, outperforming both Logistic Regression and Random Forest. Twelve clinical volunteers, comprising final-year medical students and practising clinicians, reviewed SHAP, LIME, and rule-based explanations for a shared set of de-identified sample cases, rating each technique on clarity, clinical plausibility, and trustworthiness. SHAP explanations scored highest on clarity and trustworthiness (4.3 and 4.1 out of 5), closely followed by rule-based explanations (4.0 and 4.2), while LIME scored lowest on stability, with several evaluators noting that repeated LIME runs on similar cases sometimes surfaced different top features, a known consequence of its local sampling procedure. The trained models and explanation modules were integrated into a Flask-based clinical decision-support dashboard combining a diagnostic risk score, a SHAP-based feature-contribution chart, and, where applicable, a corresponding rule, with an average combined prediction-and-explanation response time of 0.31 seconds. The study concludes that combining several complementary explanation techniques, empirically validated with clinical end users rather than assumed to work in advance, offers a more clinically grounded route to explainable medical AI than relying on a single technique in isolation.

₦5,000View
Machine Learning Intrusion Detection System: Building a Smarter Network DefenceComputer Science

Machine Learning Intrusion Detection System: Building a Smarter Network Defence

Scholarnesthub Admin

About This Research Topic Every network defence team eventually runs into the same uncomfortable truth: attackers do not always repeat themselves. A machine learning intrusion detection system is built to deal with exactly that problem. Rather than waiting for a security analyst to write a new rule for every fresh attack pattern, it learns the underlying shape of normal and malicious traffic directly from data, so that it can flag suspicious behaviour even when the exact attack has never been logged before. This article walks through a complete undergraduate research project that puts that idea to the test. It compares classical machine learning, ensemble methods, and a stacked ensemble model across two well-known network security datasets, tests three feature-selection strategies, and wraps the strongest model in a working alert dashboard. If you are exploring similar territory for your own final-year research, our Computer Science project topics library has further reference projects that show how this kind of methodology chapter, results chapter, and system implementation typically come together. Main Abstract Signature-based intrusion detection systems remain useful, but they share one structural weakness: they can only catch what they already recognise. Once an attacker varies their technique even slightly, a purely signature-driven system has no way of raising an alarm, because no matching entry exists in its database. This weakness is what has pushed so much recent research toward machine-learning-based detection, which builds a model of what normal and malicious traffic look like from historical data rather than from a fixed rulebook. This study designs, builds, and tests a machine-learning-based network intrusion detection pipeline, following the Design Science Research approach alongside the CRISP-DM process for the data-driven stages of the work. Two public benchmark datasets anchor the evaluation: NSL-KDD, a cleaned-up successor to the long-standing KDD Cup 1999 dataset, and CICIDS2017, a newer and more realistic dataset built by the Canadian Institute for Cybersecurity that captures a wider spread of modern attack behaviour, including brute-force attempts, denial-of-service traffic, web-based attacks, infiltration, and botnet activity. After cleaning and encoding the data, three feature-selection techniques were tested side by side — Recursive Feature Elimination, Mutual Information, and Lasso-based selection — ahead of training four models: Naive Bayes, Random Forest, XGBoost, and a stacked ensemble that blends Random Forest, XGBoost, and Extra-Trees under a Logistic Regression meta-model. Both a binary classification task (normal versus attack) and a multi-class task (identifying the specific attack type) were evaluated. The stacked ensemble, paired with Recursive Feature Elimination, came out on top across both datasets, reaching 99.6% accuracy and a 99.4% F1-score on CICIDS2017, and 99.1% accuracy with a 98.7% F1-score on NSL-KDD. It consistently outperformed the standalone Random Forest, XGBoost, and Naive Bayes models. The multi-class results told a more nuanced story: overall performance stayed strong, but recall dropped noticeably for rare attack categories such as infiltration and certain web-attack subtypes, a pattern directly tied to how few training examples those classes have relative to normal traffic. To close the loop between research and practice, the trained binary model was deployed behind a Flask-based monitoring dashboard that reads a simulated live traffic feed, scores each flow for intrusion risk, and raises a security alert when something looks malicious — with an average classification latency of just 8 milliseconds per flow. Taken together, the findings support stacked ensembles combined with careful feature selection as a strong, computationally realistic foundation for machine-learning-based intrusion detection, while cautioning that the near-perfect accuracy figures often reported in this field need to be read alongside dataset-specific class imbalance and cross-dataset generalisation, not treated as a promise of identical real-world performance.

₦5,000View
AI-Based Traffic Congestion Prediction SystemComputer Science

AI-Based Traffic Congestion Prediction System

Scholarnesthub Admin

About This Research Topic Most traffic apps tell you what is happening right now, which is a bit like checking the weather after you are already soaked. This project set out to build something genuinely predictive instead, a system that forecasts congestion up to several hours ahead of time, tested four different modelling approaches against each other, and packaged the winner behind a live dashboard. This article walks through how that system was built, from feature engineering on historical traffic sensor data to a head-to-head comparison of a statistical baseline, Random Forest, LSTM, and a hybrid CNN-LSTM architecture. Readers exploring related technical projects can browse the Computer Science project collection on ScholarNest for comparable studies in machine learning and systems design. What follows covers the background to AI-based traffic forecasting, the specific problem this project addresses, its objectives and research questions, the key technical terms used throughout, and closes with frequently asked questions for students and developers working on similar predictive systems. Main Abstract Traffic congestion imposes substantial economic, environmental, and quality-of-life costs on urban populations, through lost productivity, increased fuel consumption and emissions, and extended commute times, motivating sustained interest in predictive systems capable of forecasting congestion ahead of time to support proactive traffic management and route planning. This study addresses this problem by designing, implementing, and evaluating a machine learning system for short-term traffic congestion prediction that combines historical traffic sensor data with a simulated real-time data feed, moving beyond purely reactive, current-state traffic reporting toward a genuinely predictive capability. The study adopted the Design Science Research methodology combined with the Cross-Industry Standard Process for Data Mining for the data-driven components of the work. It used the publicly available Metro Interstate Traffic Volume dataset, comprising hourly traffic volume readings from a Minneapolis-St Paul interstate corridor spanning several years, combined with corresponding weather and holiday indicator features, from which a congestion-level target variable (Low, Moderate, High) was derived using volume and historical speed-relationship thresholds. Data were cleaned, engineered with cyclical time-of-day and day-of-week features and lagged historical volume features, and used to train and compare four models: a Seasonal ARIMA statistical baseline, Random Forest, a Long Short-Term Memory network, and a hybrid CNN-LSTM architecture combining convolutional feature extraction with recurrent temporal modelling. Models were evaluated on both a regression formulation, predicting continuous traffic volume using RMSE and MAE, and a complementary classification formulation, predicting discrete congestion level using accuracy, precision, recall, and F1-score. The hybrid CNN-LSTM model achieved the best performance on both formulations, with a regression RMSE of 312 vehicles per hour and a classification accuracy of 89.4% (macro F1-score of 87.6%) for one-hour-ahead prediction, outperforming the standalone LSTM (RMSE 356, accuracy 85.1%), Random Forest (RMSE 428, accuracy 79.8%), and SARIMA (RMSE 612, accuracy 68.3%). Prediction accuracy degraded gracefully with increasing forecast horizon, remaining above 80% classification accuracy up to a three-hour-ahead horizon before declining more sharply. The trained model was deployed behind a Flask-based dashboard that ingests a simulated real-time traffic feed, replayed historical data standing in for a live sensor connection, and displays current and predicted congestion levels for the monitored corridor on an interactive map-style interface, with an average end-to-end prediction latency of 45 milliseconds. The study concludes that hybrid CNN-LSTM architectures, informed by both historical patterns and current real-time conditions, offer a practical basis for short-term traffic congestion prediction capable of supporting proactive traffic management, and recommends extension to a road-network-wide, spatially-aware modelling approach as a direction for future work.

₦5,000View
A Facial Recognition Attendance System with Anti-Spoofing MeasuresComputer Science

A Facial Recognition Attendance System with Anti-Spoofing Measures

Scholarnesthub Admin

About This Research Topic A facial recognition attendance system sounds foolproof until someone simply holds up a photo to the camera. That's the gap most systems leave open, and it's exactly what this study set out to close — building a system that checks not just who is in front of the camera, but whether they're actually there, live, in the flesh. This piece walks through how a dedicated anti-spoofing stage was built, benchmarked, and combined with face recognition into a working attendance system. Readers interested in how applied machine learning performs on other real-world classification problems may also want to look at our project on predicting hospital readmission rates with machine learning , which covers a different domain but a similarly structured evaluation approach. What follows carries the full research structure — background, problem statement, aim and objectives, research questions, significance, scope, and definitions — rebuilt for a wider readership while preserving the original study's technical focus and reported results. Main Abstract Manual and card- or fingerprint-based attendance systems remain widely used in academic and workplace settings despite well-documented weaknesses: manual roll-call is time-consuming and susceptible to proxy attendance, while card- and fingerprint-based systems, though automated, still permit proxy attendance through credential sharing and raise hygiene concerns in shared-device settings. Facial recognition offers a contactless, difficult-to-share biometric alternative, but a system that only performs identity matching remains vulnerable to presentation attacks, in which an impostor presents a printed photograph, a video replay, or a mask of an enrolled individual in place of their own live face. This study addresses that problem by designing, implementing, and evaluating a facial recognition-based attendance system that incorporates an explicit anti-spoofing (liveness detection) stage, ensuring attendance is recorded only for a live, physically present individual rather than a static or replayed representation. Following the Design Science Research methodology combined with the Cross-Industry Standard Process for Data Mining, a face-recognition enrolment dataset of 40 volunteer individuals (roughly 25 images per individual, captured under varied lighting and pose) was combined with the publicly available CelebA-Spoof dataset for anti-spoofing model training, comprising live and spoof (print, replay, and cut-photo) face images. Face detection used a Multi-task Cascaded Convolutional Network, face recognition embeddings were generated using a pretrained FaceNet model, and identity matching was performed via cosine-similarity comparison against enrolled embeddings. For anti-spoofing, a MobileNetV2-based binary CNN classifier was benchmarked against a classical Local Binary Pattern texture baseline with an SVM classifier, and an eye-blink-based liveness heuristic using the Eye Aspect Ratio. The MobileNetV2 anti-spoofing model achieved the best performance — 97.8% accuracy, 97.2% precision, 98.1% recall, and a 97.6% F1-score on a held-out test set spanning print, replay, and cut-photo attacks — outperforming the LBP+SVM baseline (89.4% accuracy) and the EAR-based blink heuristic (81.7% accuracy, and specifically vulnerable to video replay attacks that include natural blinking). The face-recognition component achieved a rank-1 identification accuracy of 98.5% and a false acceptance rate of 0.6% on the 40-person enrolment set. The combined system, implemented as a desktop/web-hybrid application using OpenCV for camera capture, achieved an average end-to-end attendance-marking time of 1.1 seconds per individual. The study concludes that combining a dedicated CNN-based anti-spoofing stage with embedding-based face recognition substantially improves resistance to common presentation attacks relative to either face recognition alone or simple heuristic liveness checks, and recommends periodic model updates and expansion to 3D-mask attack resistance as directions for future work.

₦5,000View
Machine Learning for Credit Risk Scoring in Microfinance and Fintech LendingComputer Science

Machine Learning for Credit Risk Scoring in Microfinance and Fintech Lending

Scholarnesthub Admin

About This Research Topic Most credit scoring was built for people who already have a credit history — which is exactly the problem for the millions of microfinance and fintech borrowers who don't. This piece walks through a machine learning approach built specifically for that gap: blending conventional loan data with alternative signals like mobile-money activity, benchmarking several modelling approaches against each other, and — just as importantly — checking whether the resulting model treats different borrower groups fairly. Readers curious about how machine learning performs on related prediction tasks may also want to look at our project on machine learning algorithms for credit risk prediction , which covers a closely related modelling problem in more depth. What follows carries the full research structure — background, problem statement, aim and objectives, research questions, significance, scope, and definitions — rebuilt for a wider readership while preserving the original study's technical focus and reported results. Main Abstract Access to credit remains a critical enabler of small business growth and household resilience in developing economies, yet microfinance institutions and fintech lenders serving these markets frequently lack the extensive, formal credit history data that traditional credit scoring relies upon in established banking systems — a condition commonly termed the thin-file problem. This study addresses that problem by designing, implementing, and evaluating a machine learning model for credit risk scoring that combines conventional loan-application and repayment-history features with alternative, non-traditional behavioural indicators, while explicitly evaluating model performance and fairness properties relevant to responsible lending. Following the Design Science Research methodology combined with the Cross-Industry Standard Process for Data Mining, the study used two complementary datasets — the Statlog (German Credit) benchmark and a Kaggle-sourced microfinance/small-business loan dataset incorporating mobile-money transaction regularity, utility-payment history, and business-registration status — comprising 9,578 loan records after cleaning and combination. Four models were trained and compared for binary default-risk classification: Logistic Regression, Random Forest, XGBoost, and a feed-forward Artificial Neural Network, evaluated on accuracy, precision, recall, F1-score, and ROC-AUC, with particular attention to recall on the default (minority) class given the asymmetric cost of misclassifying a genuinely high-risk borrower as low-risk. Model interpretability was addressed using SHAP values, and a fairness audit compared false-positive and false-negative rates across gender and business-sector subgroups. XGBoost achieved the best overall performance — 88.7% accuracy, 79.4% default-class recall, 74.1% precision, an F1-score of 76.7%, and a ROC-AUC of 0.91 — outperforming Logistic Regression (61.3% recall), Random Forest (74.8% recall), and the ANN (76.2% recall). SHAP analysis identified prior repayment delinquency, debt-to-income ratio, and mobile-money transaction regularity as the three most influential predictors of default risk, with the alternative mobile-money feature contributing meaningfully alongside conventional financial attributes. The fairness audit found a modest but non-negligible 6.1 percentage-point disparity in false-positive rate between gender subgroups, flagged as a finding requiring further mitigation rather than a settled result. The trained model was deployed behind a Flask-based loan-officer decision-support dashboard returning a risk score, a recommended decision band, and the top SHAP-derived contributing factors per application, with an average scoring response time of 0.09 seconds. The study concludes that gradient-boosted models incorporating alternative behavioural features can meaningfully improve default-risk identification relative to conventional logistic-regression-based scorecards common in microfinance practice, while underscoring that fairness auditing and human-in-the-loop review remain essential complements to model deployment in a lending context with direct financial consequences for applicants.

₦5,000View
An AI-Powered Chatbot for Student Academic Advising and Course Registration SupportComputer Science

An AI-Powered Chatbot for Student Academic Advising and Course Registration Support

Scholarnesthub Admin

About This Research Topic Registration week has a familiar rhythm on most university campuses: long queues outside the adviser's office, the same handful of questions repeated dozens of times a day, and students with genuinely complicated situations stuck waiting behind ones who just want to confirm a prerequisite. It is a capacity problem more than a knowledge problem, and it is exactly the kind of bottleneck conversational AI is well suited to relieve. Students exploring a similarly applied AI project can browse ScholarNest's computer science project topics for related ideas in natural language processing and intelligent systems. This article walks through a complete undergraduate research project built around that exact problem: an AI-powered chatbot that handles routine academic advising and course registration queries through natural conversation, and hands off anything genuinely complex to a human adviser. Rather than settling for a single intent classifier, the study benchmarks three approaches — a classical TF-IDF/SVM baseline, a BiLSTM, and a fine-tuned DistilBERT transformer — then wraps the strongest model inside a full dialogue system grounded in a structured course and policy knowledge base, and tests it end-to-end with real users. What follows breaks down the study's background, problem statement, objectives, and scope, for students, researchers, and anyone curious about how conversational AI is being applied to student support. Main Abstract Academic advising and course registration support are essential but resource-intensive services in tertiary institutions, typically requiring students to queue for limited adviser appointments to resolve routine questions about course prerequisites, registration deadlines, credit-load limits, and graduation requirements. This demand is heavily concentrated around the start of each semester, and it frequently overwhelms available advising capacity, leaving many routine student queries unresolved in a timely manner. This study designs, implements, and evaluates an AI-powered chatbot capable of handling common academic advising and course registration queries through natural language conversation, escalating only genuinely complex or policy-ambiguous cases to a human adviser. The work follows a Design Science Research methodology paired with an Agile development approach for the conversational system itself. A corpus of 3,600 utterances, collected through a structured student survey soliciting example questions and synthetically generated paraphrases, was manually labelled across fourteen intent categories — course prerequisite inquiry, registration deadline inquiry, credit-load inquiry, GPA calculation, and adviser escalation, among others — and used to train and compare three intent classification approaches: a TF-IDF plus Support Vector Machine baseline, a Bidirectional LSTM classifier, and a fine-tuned DistilBERT transformer classifier. The chatbot's dialogue manager combines the intent classifier with a slot-filling component for extracting entities such as course codes and semesters, and a rule-based decision engine that queries a structured knowledge base of courses, prerequisites, and registration policies to generate a response. The fine-tuned DistilBERT model achieved the best intent classification performance, with an accuracy of 93.8% and a macro-averaged F1-score of 92.6%, outperforming the BiLSTM (89.1% accuracy) and the SVM baseline (83.4% accuracy). In an end-to-end task-completion evaluation involving 15 test users completing 5 representative advising tasks each, the chatbot achieved a task-completion rate of 86.7%, with most incomplete tasks attributable to queries falling outside the chatbot's trained intent set and correctly escalated to a human adviser. System testing showed an average response time of 0.6 seconds per turn. A System Usability Scale evaluation returned a mean score of 78.4, corresponding to a 'good' usability rating. The study concludes that an intent-classification-driven chatbot, grounded in a structured institutional knowledge base, can meaningfully reduce the routine advising burden on human academic advisers while reliably escalating queries beyond its competence, and recommends integration with a live student information system and periodic retraining on real deployment queries as future work.

₦5,000View
Predictive Maintenance Model for Industrial Equipment Using Sensor DataComputer Science

Predictive Maintenance Model for Industrial Equipment Using Sensor Data

Scholarnesthub Admin

About This Research Topic A machine that fails without warning does more than stop production. It forces a scramble for spare parts, idles an entire line, and often costs far more to fix under pressure than it would have under a planned schedule. For decades, manufacturers have managed that risk with two blunt tools: run equipment until it breaks, or service it on a fixed calendar regardless of its actual condition. Predictive maintenance offers a third option, using sensor data and machine learning to flag a failing component before it fails, so intervention happens only when it is genuinely needed. Students exploring a similarly applied machine learning topic can browse ScholarNest's computer science project topics for related ideas in sensor data, classification, and industrial AI. This article walks through a complete undergraduate research project built around that exact problem, tackled from two complementary angles: predicting whether a piece of equipment is about to fail, and estimating how much useful operating life a degrading component has left. The study benchmarks four models for the failure-classification task and two for the remaining-useful-life estimation task, using two independently sourced datasets, then deploys the strongest classifier behind a live monitoring dashboard. What follows breaks down the study's background, problem statement, objectives, and scope, for students, researchers, and anyone curious about how machine learning is being applied to industrial reliability. Main Abstract Unplanned downtime caused by unexpected industrial equipment failure remains one of the most significant sources of lost productivity and maintenance cost in manufacturing environments. That reality has driven a sustained shift away from purely reactive, run-to-failure maintenance and fixed-interval preventive maintenance, toward predictive maintenance, where sensor-derived condition data is used to anticipate impending failure and schedule intervention only when it is genuinely warranted. This study designs, implements, and evaluates a machine learning model for predicting industrial equipment failure from multivariate sensor data, and extends that capability to a remaining-useful-life (RUL) estimation task for a degrading component. The work follows a Design Science Research methodology paired with the CRISP-DM process for its data-driven components. Two complementary datasets were used: the AI4I 2020 Predictive Maintenance dataset, a synthetically generated but operationally realistic set of 10,000 milling-machine operating records with binary failure labels and failure-mode annotations, used for the failure-classification task; and the NASA C-MAPSS turbofan degradation dataset, used for the remaining-useful-life regression task. Data were cleaned and engineered with rolling-window statistical features (mean, standard deviation, and rate of change over sliding sensor-reading windows), then used to train and compare four models for failure classification — Logistic Regression, Random Forest, Gradient Boosting via XGBoost, and a one-dimensional Convolutional Neural Network applied to short sensor-reading sequences — and two models for RUL regression: a Random Forest Regressor and a Long Short-Term Memory (LSTM) network. XGBoost delivered the strongest failure-classification performance, reaching 98.4% accuracy, 91.7% precision, 88.3% recall, and an 89.9% F1-score on the minority failure class, ahead of Logistic Regression (71.2% F1-score), Random Forest (86.1% F1-score), and the 1D-CNN (87.4% F1-score). For RUL estimation, the LSTM model achieved a Root Mean Squared Error of 18.9 cycles, outperforming the Random Forest Regressor's 24.6 cycles. The best-performing failure-classification model was deployed behind a Flask-based monitoring dashboard that ingests simulated streaming sensor readings, displays a live equipment health status and failure-risk score for each monitored unit, and generates a maintenance alert once the risk score crosses a configurable threshold. System testing showed an average per-reading inference latency of 12 milliseconds, comfortably supporting near-real-time monitoring. The study concludes that gradient-boosted tree models offer a strong, computationally efficient basis for sensor-based failure classification on tabular condition-monitoring data, while recurrent architectures hold a meaningful advantage for sequence-dependent remaining-useful-life estimation, and recommends integration with real industrial IoT sensor streams and cost-sensitive threshold tuning as directions for future deployment.

₦5,000View

Can't find your topic? Request a custom material →