- Background subtraction: This is a technique to separate moving objects from the static background of a scene. It works by comparing each frame of the video to a background pattern and spotting the differences. These differences are considered to be part of the foreground.
- Shadow detection and morphological operations: After identifying the foreground, shadow detection helps distinguish shadows from real objects. Morphological operations such as erosion and dilation are applied to improve image quality by removing noise and strengthening the structure of detected objects.
- Histogram of Oriented Gradients (HOG): HOG extracts features from images by computing directional gradients in various portions of the image. These features are then used to identify specific shapes, such as the outline of a person.
- Support Vector Machines (SVM): SVM is a classification technique that uses a training data set to determine a hyperplane that best separates different classes. In the case of human detection, SVM classifies the data based on the features extracted by HOG as human or non-human.
- Appearance-based methods: These methods focus on the appearance and shape of objects to track them over time. Tracking is based on comparing the appearance of the object in consecutive frames, thus maintaining coherent tracking.
vineri, 12 ianuarie 2024
Human detection and tracking systems using OpenCV
marți, 12 decembrie 2023
A brief introduction to MLOps
What is MLOps?
Machine Learning operations involve applying a set of processes, or more specifically, a sequence of steps to integrate a Machine Learning model into the production environment. There are several stages to go through before such a model is ready for deployment, and these processes ensure that the model can be scaled to meet a large user base and provide accurate results.
Why do we need MLOps?
DevOps vs MLOps
MLOps Workflow
#1 Data preparation and management
The first phase in any machine learning process is the collection of data. Without clean and accurate data, the models are useless. Hence, data preparation & management is a crucial phase in MLOps workflow.It involves collecting, cleaning, transforming, and managing data so that we have the correct data in place to train the ML models. The end goal is to have data that is complete and accurate.
#2 Model training and validation
#3 Model deployment
After the model is trained and the performance is validated, it’s time to deploy it to production. The model deployment phase involves taking the model that has been trained, tested, and validated in earlier phases and making it available for use by other applications or systems.
#4 Continuous model monitoring & retraining
Machine learning models aren’t something like deploy and forget. One needs to constantly monitor the model’s performance and accuracy. Since the real world can change which might affect the efficiency and accuracy of the model.
Popular tools:
luni, 11 decembrie 2023
Cirrhosis Patient Survival Prediction
Descrierea studiului:
Setul de date "Cirrhosis Patient Survival Prediction" este o colecție de informații care vizează predicția supraviețuirii pacienților cu ciroză hepatică. Acesta cuprinde 17 caracteristici și variabile pentru a prezice supraviețuirea și starea pacienților diagnosticați cu ciroză hepatică, inclusiv date demografice, rezultate ale testelor de laborator, informații despre tratamente și alte factori medicali relevanți.
Supraviețuirea este codificată 0 = D – deces, 1 = C – cenzurat, 2 = CL – cenzurat datorită transplantului
Obiectivul principal al acestui set de date este de a permite analiza și predicția supraviețuirii pacienților în funcție de diferitele lor caracteristici și factori medicali. Ciroza este consecința afectării prelungite a ficatului, care duce la cicatrici extinse, adesea din cauza unor afecțiuni precum hepatita sau consumul cronic de alcool. Datele sunt furnizate dintr-un studiu clinic al Clinicii Mayo privind ciroza biliară primară (CBP) a ficatului (1974 și 1984).
Înțelegerea setului de date:
Setul de date poate conține informații despre diversitatea răspunsurilor pacienților la tratamentele specifice, evoluția bolii și alte detalii clinice care pot fi cruciale pentru înțelegerea și gestionarea cirozei hepatice.
424 de pacienți cu PBC care s-au prezentat la Clinica Mayo s-au calificat pentru un studiu randomizat controlat cu placebo care a testat medicamentul D-penicilamină. Dintre aceștia, primii 312 pacienți au luat parte la studiu și au în mare parte date cuprinzătoare. Restul de 112 pacienți nu s-au alăturat studiului clinic, dar au fost de acord să înregistreze valorile de bază și să fie supuși urmăririi supraviețuirii. Șase dintre acești pacienți au fost în scurt timp imposibil de urmărit după diagnosticul lor, lăsând date pentru 106 dintre acești indivizi, în plus față de cei 312 care au făcut parte din studiul randomizat.
Acest lucru poate fi utilizat pentru a dezvolta modele predictive care să ajute medicii să evalueze riscurile și perspectivele de supraviețuire pentru pacienții diagnosticați cu ciroză hepatică. Folosind datele din acest set, algoritmi de învățare automată pot fi antrenați pentru a identifica tipare și corelații între variabilele din cadrul acestui context medical specific.
Este important să se menționeze că manipularea și analiza acestui set de date trebuie realizată cu mare atenție și etică medicală, respectând confidențialitatea și drepturile pacienților, precum și utilizând informațiile doar în scopuri de cercetare și îmbunătățire a îngrijirii medicale.
Variabilele setului de date:
Procesarea datelor și cercetarea
https://colab.research.google.com/drive/1752vdy9Zb0MLbEoB-veg5QvRLuClMdgc#scrollTo=dceff5d2
Pentru început am împărțit setul de date în următoarele capitole:
Încărcarea bibliotecilor si a setului de date
Informații despre setul de date
Validarea
Distribuția caracteristicilor numerice
Distribuția caracteristicilor categorice
Distribuția țintita
Colorarea și gruparea ierarhica
Pregătirea setului de date
Validarea încrucișata a modelului
Predicția și transmiterea rezultatelor
Am încărcat bibliotecile: numpy, pandas, matplotlib.pyplot și seaborn, împreună cu module din sklearn, scipy și module din bibliotecile de clasificare precum: xgboost, lightgbm, catboost. Apoi am încărcat seturile de date de antrenament, de test și setul de date original.
Am studiat prin statistică descriptivă toate cele 3 seturi de date.
Am efectuat validarea contradictorie pentru a vedea dacă seturile de antrenament și cel de test au distribuție similara prin determinarea scorului ROC-AUC, cu un rezultat de 0,50192, trăgând concluzia că cele două seturi sunt similare.
Am suprapus în grafice distincte pentru fiecare caracteristică din tabele (coloană) distribuția valorilor numerice observând că cele două seturi de antrenament și de test au o distribuție similară, iar setul de date original o distribuție aparte (albastru).
Am aplicat același lucru pentru variabilele categorice pentru datele din setul de antrenament, creându-ne o imagine despre câți pacienți au primit D-penicilamină, câți placebo, despre distribuția pe sexe (majoritatea fiind bărbați – 93%), despre proporția prezenței ascitei (5%), apariția hepatomegaliei a jumătate dintre pacienți.
Distribuția țintită a pacienților în setul de date de antrenament relevă că 63% au supraviețuit, 34% au decedat, iar 3% au supraviețuit în urma transplantului hepatic.
Am realizat corelarea și gruparea ierarhică prin realizarea corelațiilor dintre caracteristicile seturilor de date sub forma unui heatmap, prin care s-a observat o corelație puterinică între nivelurile serice ale cuprului și bilirubinei, al SGOT (aspartat aminotransferazei - AST) cu bilirubina, nivelul de cupru și al fosfatazei alcaline.
Am realizat pregătirea setului de date pentru Machine Learning models și am generat modele precum: regresie logistică, analiza discriminarii liniare, distribuție Gaussiană, distribuție Bernoulli, clasificarea K-Neighbors, Random forest, XGBC, LGBMC, catboost etc.
Evaluarea rezultatelor
Rezultatele generate de modelele noastre pentru regresii, distribuții și clasificări le-am reprezentat grafic într-un barplot care ne sugerează că cel mai bun model pentru următoarele predicții este GradientBoostingClassifier. În urma rezultatelor antrenăm modelul pe setul de date de antrenament, urmând a face predicțiile.
Perspective de viitor
În viitor propunem să realizăm predicțiile pentru aceste seturi de date.
Bibliografie
https://www.kaggle.com/datasets/joebeachcapital/cirrhosis-patient-survival-prediction/data
vineri, 8 decembrie 2023
Artificial Intelligence approaches to Chatbot Development
Artificial Intelligence approaches to Chatbot Development
INTRODUCTION :
STATE OF THE ART :
IMPORTANCE :
- 24/7 Availability and Instant Responses
- Time and Cost Savings
- Improved User Experience
- Task automation
- Ability to speak multiple languages
HOW DO THEY WORK :
CONCLUSION :
REFERENCES :
miercuri, 22 noiembrie 2023
Machine learning in software defect prediction
Introduction
In recent times, there has been a substantial increase in the quantity, scale, and intricacy of software systems. These significant developments have heightened the need for software testing, a process that is both resource-intensive and time-consuming[1]. Software Defect Prediction (SDP) is indeed crucial for identifying potentially defective software modules early in the development process. To optimize resource allocation and minimize testing costs, it is important not only to identify defective modules but also to prioritize them effectively.
Ensuring the reliability of software is a paramount objective, and in this pursuit, Software Quality Assurance (SQA) teams assume a pivotal role within the software development process. Consequently, the strategic prioritization of SQA activities emerges as a crucial phase in the SQA lifecycle. A fundamental aspect of this prioritization involves the application of Software Defect Prediction (SDP) methodologies, which serve the purpose of identifying high-risk software components and assessing the impact of various software metrics on the probability of failure in software modules. The perpetual quest for more advanced and refined SDP models underscores the ongoing necessity for sophisticated tools and methodologies in this realm.
The predictive process for identifying software modules with defect proneness, commonly known as Software Defect Prediction (SDP), is a comprehensive approach aimed at assessing the likelihood of bugs or defects in various modules based on their method-level and class-level metrics. This method involves utilizing historical data and statistical models to predict which modules are more likely to have issues, allowing for a proactive and strategic allocation of resources during the testing phase.[2]
In essence, SDP goes beyond mere bug detection during testing; it is a proactive strategy that helps software development teams prioritize their testing efforts more effectively. By analyzing the characteristics of software modules at both the method and class levels, development teams can gain insights into potential vulnerabilities or areas of concern. This predictive analysis aids in early identification of modules that may be more susceptible to defects, enabling teams to focus their testing efforts where they are needed most.
How it works
Data Collection and Feature Extraction:Machine learning models require data for training. In the context of SDP, historical data related to software development, including defect information, is collected. Features, representing various characteristics of software modules (e.g., code complexity, size, historical defect data), are extracted from this dataset.
Training the Model:Supervised learning algorithms, such as Decision Trees, Random Forests, Support Vector Machines, or Neural Networks, are commonly employed. The model is trained on the historical dataset, learning patterns and relationships between the extracted features and the occurrence of defects.
Cross-Validation:To ensure the model's generalizability and robustness, cross-validation techniques are often employed. This involves splitting the dataset into multiple subsets, training the model on some subsets, and validating its performance on the remaining subsets.
Feature Importance Analysis:ML models allow for the analysis of feature importance, indicating which features contribute more significantly to the prediction of defects. This analysis can provide insights into the factors that make certain software modules more defect-prone.
Handling Imbalanced Data:Since software defect datasets are often imbalanced (few modules have defects compared to the total), ML models need techniques to handle this imbalance. Sampling methods and specialized algorithms are employed to address this issue.
Continuous Improvement:ML models can continuously learn and adapt as new data becomes available. This enables the SDP system to evolve and improve its predictive capabilities over time.
vineri, 17 noiembrie 2023
The Impact of Generative AI and Large Language Models
Generative AI and Large Language Models (LLMs) have become game
changers in artificial intelligence. These advanced systems, powered by complex
algorithms and extensive datasets, are pushing the limits of what machines can
achieve: creativity, problem-solving, and human-like interaction.
The Current State of Generative AI
State-of-the-art generative AI models such as OpenAI's GPT-4
("the latest step in OpenAI's effort to scale deep learning") possess
an unprecedented ability to generate human-like text and more, making them
valuable tools for multiple applications. For example, GPT-4 passes a simulated
bar exam with a score around the top 10% of test takers; in contrast, GPT-3.5's
score was around the bottom 10%.
Why Generative AI Matters
Due to its ability to improve various processes, Generative AI
is of great importance. The applications are vast and, in many forms, from
content creation and language translation to code generation and creative
writing. The ability to generate coherent and contextually relevant text
empowers businesses and individuals alike, delivering a new level of efficiency
and innovation.
Cloud Resources and Solutions
Cloud solutions play a base role in making powerful generative
AI accessible to the audience. By utilizing the scalability and computing
resources of cloud platforms, users can take advantage of the capabilities of
these models without the need for extensive hardware infrastructure.
In addition, cloud providers come with a wide variety of ready-to-use integrated AI resources. Cloud providers have announced a wide range of resources that are already available (for example, many APIs can be used to create chatbots, virtual assistants, and more).
Examples
The City of Kelowna uses AI technology, specifically Azure
OpenAI Service and Azure Cognitive Search, to develop an intelligent search
solution for public services. This system addresses citizen requests using the
available information and ensures strict compliance with data privacy measures.
Generative AI-powered chatbots and virtual assistants provide
fast and accurate answers to customer questions, providing personalized
recommendations and assistance. It improves overall customer service, reduces
wait times, improves operational efficiency, and increases satisfaction. For
example, Azure Bot Services enables the easy creation of bots for non-technical
people and significantly reduces time and costs.
Future of Generative AI
Generative AI and LLM have opened new doors in the world of AI,
pushing the boundaries of what machines can achieve. As we move forward and
integrate these technologies across different domains, using their
advancements, we reshape how we interact with and benefit from artificial
intelligence.
In conclusion, the future is bright for generative AI, with
continued research and development for even more sophisticated models and
applications. As these technologies evolve, their impact on industries and
everyday life is high.
References:
https://www.techtarget.com/searchenterpriseai/definition/generative-AI
https://openai.com/research/gpt-4
https://azure.microsoft.com/en-us/blog/welcoming-the-generative-ai-era-with-microsoft-azure/
marți, 14 noiembrie 2023
Unveiling the Future: Facial Recognition Tech as a Health Sentinel
Greetings, fellow tech enthusiasts! Today, we embark on a journey into the cutting-edge realm of facial recognition technologies, where pixels meet emotions, and algorithms decipher the intricate language of the human face. Buckle up, because the future is here, and it's brimming with possibilities.
Premise: Decoding Emotions in Pixels
Picture this: recent strides in emotion recognition systems, showcased in [1], have thrust artificial intelligence into the spotlight, enabling it to unravel the subtle nuances of human emotions through facial expressions. It's not just about recognising a smile or a frown; it's about understanding the complex dance of emotions painted on our faces.
Advancements in Health Detection ([2]): A Glimpse into Tomorrow
Now, let's fast forward to [2], where the plot thickens. The hypothesis takes a bold turn, suggesting that this facial emotion recognition technology isn't merely a spectator of human sentiments but a potential game-changer in predicting psychiatric illnesses and latent mental health issues. Our faces might just hold the key to unlocking the mysteries of our minds.
Challenges in the Facial Recognition Frontier ([3]): Illuminating the Path Ahead
Of course, no epic journey is without its challenges. [3] sheds light on the obstacles in the facial expression recognition (FER) quest. From battling illumination issues to navigating the maze of occlusions, the road ahead is complex. Yet, in these challenges lie the seeds of opportunities.
Peering into the Health Horizon ([3]): Detecting Ailments through Emotions
Despite the hurdles, the study suggests that automated emotion detection is not just a tech marvel but a potential health sentinel. The whispers of our emotions might just serve as early indicators, pointing towards a myriad of health conditions. Imagine a world where your face not only mirrors your feelings but also signals potential health concerns.
The Deep Dive into Facial Sentiment Analysis ([3]): A Geek's Delight
Now, let's geek out with [3], a comprehensive survey that dissects the state-of-the-art machine learning and deep learning approaches in Facial Sentiment Analysis. It's not just about recognising a smile; it's about the algorithms that dance through pixels, unraveling the science behind the sentiment.
Conclusion: Bridging Pixels and Health, A New Frontier Unveiled
In the grand finale, we find ourselves at the crossroads of pixels and health. Facial recognition technologies aren't just transforming the way we decode emotions; they're opening doors to a future where our faces become gateways to understanding both the mind and the body. As the algorithms evolve, so does our comprehension of the intricate tapestry of human existence.
References:
[1]: https://ieeexplore.ieee.org/abstract/document/9091188
[2]: https://dl.acm.org/doi/abs/10.1145/3474124.3474205
[3]: https://pdfs.semanticscholar.org/f6c5/777623dcfc7d2cd74aa0957791100aca8b67.pdf
Gestionarea traficului prin inteligenta artificiala
Gestionarea traficului Circulația rutieră devine din ce în ce mai aglomerată și mai lentă, favorizându-se producerea a numeroa...
-
Introduction In today's world of rapidly advancing technology, the widespread growth of digital environments has created both opportun...
-
In the context of increasing security threats such as terrorism, advanced video surveillance systems are becoming essential. These systems n...
-
The importance of fraud detection in banking Fraud detection and prevention is a crucial aspect of today’s financial industry as it preven...


