Scientific Journal of Astana IT University
Not a member yet
    250 research outputs found

    EFFECTIVENESS OF MACHINE LEARNING METHODS IN DETERMINING EARTHQUAKE PROBABLE AREAS: EXAMPLE OF KAZAKHSTAN

    Get PDF
    This study investigates the effectiveness of machine learning methods in identifying earthquake-prone areas in Kazakhstan and its neighboring regions. By leveraging a comprehensive dataset encompassing significant earthquake data from 1900 to 2023, various machine learning algorithms were employed, including RandomForest, GradientBoosting, Logistic Regression, Support Vector Classification (SVC), K-Nearest Neighbors (KNeighbors), Decision Tree, XGBoost, LightGBM, AdaBoost, and MLPClassifier. The primary objective was to analyze and compare the performance of these models in predicting earthquake magnitudes and frequencies. The results reveal that certain algorithms significantly outperformed others in terms of accuracy, underscoring the potential of machine learning techniques to enhance earthquake prediction capabilities. Notably, XGBoost and RandomForest demonstrated the highest predictive accuracy, suggesting their suitability for application in seismic risk assessment. These findings offer valuable insights for governmental agencies engaged in disaster management and prevention planning, highlighting the practical implications of integrating advanced analytical techniques in their strategies. In addition to model performance analysis, a visual heatmap was generated to illustrate the geographical distribution of earthquake occurrences across the studied regions. This visual representation effectively identifies high-risk areas, serving as a crucial tool for local authorities and researchers in making informed decisions regarding safety measures and emergency preparedness. This research contributes to the expanding body of knowledge on earthquake prediction utilizing machine learning, emphasizing the necessity for continuous improvement in predictive models by incorporating additional environmental and geological factors. The implications of these findings extend beyond academic discourse, holding significant potential for enhancing public safety in regions vulnerable to seismic activity. As such, this study advocates for the integration of machine learning methodologies in disaster management frameworks to mitigate risks and enhance preparedness in earthquake-prone regions

    DEVELOPMENT OF MACHINE LEARNING METHODS FOR MARKET TRENDS

    Get PDF
    In the rapidly evolving real estate market, the application of machine learning (ML) is crucial for understanding and predicting price trends. This study evaluates and compares seven ML models, including multiple linear regression, random forest regression, support vector regression (SVR), decision tree regression, and XGBoost, to determine the most effective predictor of real estate prices in Astana, Kazakhstan. The study focuses on the Yesil district, a key area in the city, utilizing a dataset of over 9,000 records extracted from a broader collection of more than 30,000 real estate transactions across Kazakhstan. Through rigorous experimentation, model performance was assessed using statistical metrics such as mean absolute error (MAE), root mean square error (RMSE), and the coefficient of determination (R-squared). The results indicate that the Random Forest Regressor and XGBRegressor models outperformed others, achieving the highest R-squared values (99.55% and 99.18%, respectively) and the lowest MAE and RMSE values. These findings highlight their robustness in predicting housing prices with high accuracy. The primary objective of this study was to develop a precise ML model capable of accurately forecasting real estate prices in Astana based on key market attributes. The superior predictive performance of the Random Forest and XGBRegressor models justifies their selection for deployment in real-world applications. Their high predictive accuracy suggests their potential utility for real estate professionals, policymakers, and investors seeking data-driven insights into market dynamics. This research expands knowledge on the applications of ML in the real estate sector, reinforcing the importance of evidence-based decision-making within the industry

    THE IMPACT OF AI AND PEER FEEDBACK ON RESEARCH WRITING SKILLS: A STUDY USING THE CGSCHOLAR PLATFORM AMONG KAZAKHSTANI SCHOLARS

    Get PDF
    This research studies the impact of AI and Peer feedback on the academic writing development of Kazakhstani scholars using the CGScholar platform − the product of cutting-edge research and development into collaborative learning, big data, and artificial intelligence developed by educators and computer scientists at the University of Illinois Urbana-Champaign (UIUC). The study aimed to find out how familiarity with AI tools and peer feedback processes affects participants’ openness to incorporating feedback into their academic writing. The study involved 36 Bolashak scholars enrolled in a scientific internship focused on education at the University of UIUC. A survey with 15 questions with multiple-choice, Likert scale, and open-ended questions was employed to collect a data. The survey was conducted via Google Forms in both English and Russian to ensure linguistic accessibility. Demographic information such as age, gender, and first language were collected to provide a nuanced understanding of the data. The analysis revealed a moderate positive correlation between familiarity with AI tools and openness to making changes based on feedback, and a strong positive correlation between research writing experience and expectations of peer feedback, especially in the area of research methodology. These results show that participants are open minded to AI-assisted feedback, however they still highly appreciate peer input, especially regarding methodological guidance. This study demonstrates the potential benefits of integrating AI tools with traditional feedback mechanisms to improve research writing quality in academic settings. Further research is recommended to evaluate the long-term impact of AI and peer feedback on academic writing skills, particularly through longitudinal studies that assess skill retention over multiple feedback cycles. Additionally, expanding the study to include a more diverse academic audience will provide deeper insights into how feedback mechanisms function across different research cultures and disciplines

    USING MLOPS FOR DEPLOYMENT OF OPINION MINING MODEL AS A SERVICE FOR SMART CITY APPLICATIONS

    Get PDF
    This paper presents the MLOps strategy, which adapts the automation principles of DevOps to the deployment and lifecycle management of artificial intelligence (AI) models. By leveraging high-performance automation, MLOps ensures seamless AI development and operations integration, enabling efficient and reliable model deployment. The study demonstrates this approach by implementing the Astana Opinion Mining macro-service customized for sentiment analysis. This macro-service evaluates public opinions based on a criteria taxonomy for assessing the urban environment's sustainable development. As a smart city application, the system facilitates the collection and analysis of citizen feedback to assess the performance of city services and inform urban planning decisions. Technologically, the MLOps strategy employs containers and microservices to construct robust data and process pipelines. Four core pipelines were developed in this research: data collection, feature engineering, experimentation, deployment, and maintenance. The data collection pipeline is achieved through automated crawling from diverse sources such as social media and other internet platforms. The feature engineering pipeline ensures data preprocessing by removing noise, identifying message languages, categorizing topics, and preparing data for further analysis. The experimentation pipeline incorporates services for data labeling, model training, and performance evaluation customized to sentiment analysis tasks. Finally, the deployment pipeline and maintenance pipeline deliver trained models to end-users, ensuring their continual improvement and adaptation. Using this MLOps framework, four models of sentiment analysis were tested in Russian: "Blanchefort," "Sismetanin," "MonoHime," and "Dostoevsky." The "Blanchefort" showed an accuracy of 71,43%. The resulting MLOps framework is fault-tolerant, scalable, and enables real-time urban environment assessments. By automating workflows, the architecture enhances operational efficiency, offering practical applications for smart city initiatives and sustainable urban development, contributing to better decision-making

    DESIGN AND DEVELOPMENT OF CIRCULARLY POLARISED ANTENNA FOR RFID SYSTEM

    Get PDF
    This paper presents the research and development results of circularly polarised antennas used in radio frequency identification (RFID) systems. Such antennas play a crucial role in improving reliability, orientation independence and reading range of RFID systems in industry, transport and logistics. The frequency range under consideration is 860-900 MHz (UHF), which is widely used for passive RFID technologies due to its favorable propagation characteristics and compatibility with international standards. A printed antenna from FEIG ELECTRONIC GmbH (Germany) was used as a reference for the developed antenna. This paper presents the results of a similar antenna but without the use of a symmetry transformer. The elimination of this component reduces the overall design complexity, improves manufacturability and minimizes manufacturing cost, making the design more suitable for mass deployment. The printed dipole was developed on a 1.6 mm thick FR4 substrate with a relative dielectric constant of 4.3 and a dielectric loss tangent of 0.02. The dimensions of the developed printed dipole correspond to 332 mm × 34 mm × 1.6 mm. The printed dipole and the overall design of the developed RFID antenna were pre-simulated in the software environment “CST Studio Suite”, which allows accurate simulation of the electromagnetic behavior. This modelling step was necessary to optimize the input matching, radiation efficiency and circular polarization characteristics. The frequency of the designed antenna was 868 MHz (|S11| < -10 dB). and the radiated power was measured to be -11.7 dBm. The layout of the printed dipole was designed using Altium Designer software. The prototype assembly proceeded following model-based and electromagnetic simulation techniques. A Spectrum Rider FPH spectrum analyzer conducted test measurements which supported the theoretical prediction results. The proposed framework demonstrates great promise as an inexpensive solution with high detection efficiency for modern RFID systems operating in diverse conditions

    CHALLENGES IN GENERALIZING BREAST MRI TUMOR SEGMENTATION ACROSS MULTIPLE DATASETS

    Get PDF
    Accurate segmentation of breast tumors in dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for precise diagnosis, treatment planning, and quantitative analysis. While deep learning methods have achieved strong performance in controlled research settings, their ability to generalize across diverse clinical datasets remains underexplored and poses a major barrier to clinical adoption. In this study, we evaluate the cross-dataset generalizability of a 3D Residual U-Net model using the multicenter MAMA-MIA benchmark, which consolidates four publicly available breast MRI collections annotated by expert radiologists. A leave-one-out experimental design is employed, with three datasets used for training and validation, and the remaining dataset held-out for independent testing to simulate real-world deployment scenarios. Model performance is assessed using Dice coefficient, Precision, and Recall, alongside quantitative analysis of tumor volume estimation accuracy. The best Dice score achieved by our model was 0.683 when tested on the NACT subset. Results show a consistent degradation in segmentation accuracy when models are applied to unseen datasets, indicating that performance declines significantly outside the distribution of the training data. The most pronounced drop occurs when the DUKE dataset serves as the held-out test set, where the model struggles to adapt to differences in pre-release preprocessing strategies. A targeted qualitative review of 160 representative scans further reveals key factors contributing to both successful and failed segmentations, including variations in image field of view, temporal enhancement patterns, acquisition era, and artifact prevalence. Overall, these findings underscore the importance of accounting for dataset heterogeneity, domain shift, and standardized preprocessing in the development of robust, clinically deployable breast MRI segmentation models capable of generalizing across institutions and imaging protocols

    EKMGS: A HYBRID CLASS BALANCING METHOD FOR MEDICAL DATA PROCESSING

    Get PDF
    The field of medicine is witnessing rapid development of AI, highlighting the importance of proper data processing. However, when working with medical data, there is a problem of class imbalance, where the amount of data about healthy patients significantly exceeds the amount of data about sick ones. This leads to incorrect classification of the minority class, resulting in inefficient operation of machine learning algorithms. In this study, a hybrid method was developed to address the problem of class imbalance, combining oversampling (GenSMOTE) and undersampling (ENN) algorithms. GenSMOTE used frequency oversampling optimization based on a genetic algorithm, selecting the optimal value using a fitness function. The next stage implemented an ensemble method based on stacking, consisting of three base (k-NN, SVM, LR) and one meta-model (Decision Tree). The hyperparameters of the meta-model were optimized using the GridSearchCV algorithm. During the study, datasets on diabetes, liver diseases, and brain glioma were used. The developed hybrid class balancing method significantly improved the quality of the model: the F1-score increased by 10-75%, and accuracy by 5-30%. Each stage of the hybrid algorithm was visualized using a nonlinear UMAP algorithm. The ensemble method based on stacking, in combination with the hybrid class balancing method, demonstrated high efficiency in solving classification tasks in medicine. This approach can be applied for diagnosing various diseases, which will increase the accuracy and reliability of forecasts. It is planned to expand the application of this approach to large volumes of data and improve the oversampling algorithm using additional capabilities of the genetic algorithm

    APPLYING MACHINE LEARNING FOR ANALYSIS AND FORECASTING OF AGRICULTURAL CROP YIELDS

    Get PDF
    Analysis and improvement of crop productivity is one of the most important areas in precision agriculture in the world, including Kazakhstan. In the context of Kazakhstan, agriculture plays a pivotal role in the economy and sustenance of its population. Accurate forecasting of agricultural yields, therefore, becomes paramount in ensuring food security, optimizing resource utilization, and planning for adverse climatic conditions. In-depth analysis and high-quality forecasts can be achieved using machine learning tools. This paper embarks on a critical journey to unravel the intricate relationship between weather conditions and agricultural outputs. Utilizing extensive datasets covering a period from 1990 to 2023, the project aims to deploy advanced data analytics and machine learning techniques to enhance the accuracy and predictability of agricultural yield forecasts. At the heart of this endeavor lies the challenge of integrating and analyzing two distinct types of datasets: historical agricultural yield data and detailed daily weather records of North Kazakhstan for 1990-2023. The intricate task involves not only understanding the patterns within each dataset but also deciphering the complex interactions between them. Our primary objective is to develop models that can accurately predict crop yields based on various weather parameters, a crucial aspect for effective agricultural planning and resource allocation. Using the capabilities of statistical and mathematical analysis in machine learning, a Time series analysis of the main weather factors supposedly affecting crop yields was carried out and a correlation matrix between the factors and crops was demonstrated and analyzed. The study evaluated regression metrics such as Root Mean Squared Error (RMSE) and R2 for Random Forest, Decision Tree, Support Vector Machine (SVM) algorithms. The results indicated that Random Forest generally outperformed the Decision Tree and SVM in terms of predictive accuracy for potato yield forecasting in North Kazakhstan Region. Random Forest Regressor showed the best performance with an R2 =0.97865. The RMSE values ranged from 0.25 to 0.46, indicating relatively low error rates, and the R2 values were generally positive, indicating a good fit of the model to the data. This paper seeks to address these needs by providing insights and predictive models that can guide farmers, policymakers, and stakeholders in making informed decisions

    COMPARATIVE ANALYSIS OF FEDERATED MACHINE LEARNING ALGORITHMS

    Get PDF
    In this paper, the authors propose a new machine learning paradigm, federated machine learning. This method produces accurate predictions without revealing private data. It requires less network traffic, reduces communication costs and enables private learning from device to device. Federated machine learning helps to build models and further the models are moved to the device. Applications are particularly prevalent in healthcare, finance, retail, etc., as regulations make it difficult to share sensitive information. Note that this method creates an opportunity to build models with huge amounts of data by combining multiple databases and devices. There are many algorithms available in this area of machine learning and new ones are constantly being created. Our paper presents a comparative analysis of algorithms: FedAdam, FedYogi and FedSparse. But we need to keep in mind that FedAvg is at the core of many federated machine learning algorithms. Data testing was conducted using the Flower and Kaggle platforms with the above algorithms. Federated machine learning technology is usable in smartphones and other devices where it can create accurate predictions without revealing raw personal data. In organizations, it can reduce network load and enable private learning between devices. Federated machine learning can help develop models for the Internet of Things that adapt to changes in the system while protecting user privacy. And it is also used to develop an AI model to meet the risk requirements of leaking client's personal data. The main aspects to consider are privacy and security of the data, the choice of the client to whom the algorithm itself will be directed to process the data, communication costs as well as its quality, and the platform for model aggregation

    OPTIMIZING PROCESSOR WORKLOADS AND SYSTEM EFFICIENCY THROUGH GAME-THEORETIC MODELS IN DISTRIBUTED SYSTEMS

    Get PDF
    The primary goal of this research is to examine how different strategic behaviors adopted by processors affect the workload management and overall efficiency of the system. Specifically, the study focuses on the attainment of a pure strategy Nash Equilibrium and explores its implications on system performance. In this context, Nash Equilibrium is considered as a state where no player has anything to gain by changing only their own strategy unilaterally, suggesting a stable, yet not necessarily optimal, configuration under strategic interactions. The paper rigorously develops a formal mathematical model and employs extensive simulations to validate the theoretical findings, thus ensuring the reliability of the proposed model. Additionally, adaptive algorithms for dynamic task allocation are proposed, aimed at enhancing system flexibility and efficiency in real-time processing environments. Key results from this study highlight that while Nash Equilibrium fosters stability within the system, the adoption of optimal cooperative strategies significantly improves operational efficiency and minimizes transaction costs. These findings are illustrated through detailed 3D plots and tabulated results, which provide a detailed examination of how strategic decisions influence system performance under varying conditions, such as fluctuating system loads and migration costs. The analysis also examines the balance between individual processor job satisfaction and overall system performance, highlighting the effect of rigid task reallocation frameworks. Through this study, the paper not only improves our understanding of strategic interactions within computational systems but also provides key ideas that could guide the development of more efficient computational frameworks for various applications

    235

    full texts

    250

    metadata records
    Updated in last 30 days.
    Scientific Journal of Astana IT University is based in Kazakhstan
    Access Repository Dashboard
    Do you manage Scientific Journal of Astana IT University? Access insider analytics, issue reports and manage access to outputs from your repository in the CORE Repository Dashboard!