Multilingual Opinion mining and its stance detection by using Machine learning
Natural language Processing is referred to as a modern way of computing semantic, analyzing, and classify particular information about any textual data or Source data. Opinion mining was introduced to analyzing the data by the machine rather than the human effort. As companies and ecommerce sites sp
2025-06-28 16:28:38 - Adil Khan
Multilingual Opinion mining and its stance detection by using Machine learning
Project Area of Specialization Artificial IntelligenceProject SummaryNatural language Processing is referred to as a modern way of computing semantic, analyzing, and classify particular information about any textual data or Source data. Opinion mining was introduced to analyzing the data by the machine rather than the human effort. As companies and ecommerce sites spending a lot of money to achieve best for customers through many online feedback, or questionnaire, but is not efficient way to handle this scenario. This is a modern technique to classify the efficiency, attitude, opinions, or behavior of readers or listeners toward any product, politician, etc. In this modern era of Artificial Intelligence and Machine Learning Natural Language processing is considered to be the most prominent field of Artificial Intelligence and machine learning domain to provide the outlook and future use of our project. opinion mining has 5 different phases in changing the direction of progress and it will perform lexical analysis on scrapped data and identify the polarity of the review both it’s positive, negative or neutral, and this process Cycle includes five major steps to perform generate accurate results. By implementing this on people’s reviews or comments as raw data to generate the best opinion about products, politicians, cricketers, etc. This system uses opinion mining to track the emotions, attitudes, behaviors, and feelings of human beings towards a particular product, services, or topic e.g in Multilingual way by using different machine learning classifier.
Project ObjectivesThe key requirement & objective of this project is to create a model that detects the multilingual language opinion mining and stance detection by usng different machine learning classifier.
Project Implementation MethodThe work on Language opinion and its sentiment analysis has been very limited in Urdu and Roman Urdu languages. Roman Urdu is frequently written on blogging websites, chat rooms and advertisements and by users in the comment sections ,as mentioned earlier, there is no publicly available dataset available for such purposes. Therefore, a dataset is prepared by collecting opinions/reviews in Roman Urdu from different websites. The proposed model is divided into Four main Module. Firstly, three language data Urdu, English and roman is automatic identified from different blog. Like The users are easy to feel to post their comments in their native language that is easy to understand to other people. for example, ‘‘ye cam achi condition me hay”, ‘‘ye mobile nice hay”, ‘‘ye mobile kay feature ktnay achay he”,‘‘bcahy aj kal bht zyada mobile istamal kartay he”, ‘‘ye ek achi adat nae hain”, ‘‘is seun ke parwarish pe bura asar parta he”. This large training sentiment corpus was labeled prior to training. If an opinion contains more than 25% English word that review does include in the opinion English dataset because this research deal with English, Urdu and Roman Urdu comment. Then after that in preprocessing step, remove stop words that is unnecessary, numerical characters, punctuation marks. Words which are non-semantic in nature are termed as stop-words and usually include articles, prepositions, pronouns and prepositions. Roman stops words like ‘houn', 'ab', 'thy', kch''hain', 'han', 'to', 'is', 'thay' etc are removed. As they hold no or very little information about the opinion of the review, so they are removed from the dataset This research only allowed Roman ,Urdu character no special characters are allowed and all text are converted to lower case and stop words of roman Urdu ,English and Urdu are deleted.. Figure 2 show the step in the language classification.

Figure 2: Language Classification
Benefits of the ProjectOver 64 million people worldwide speak Urdu language and are aware of its Roman script, therefore, this study particularly focuses on multilingual low resource language processing.
Technical Details of Final DeliverableThis project will compose of Software System that give a detection model based on the language given to the system.
Case 1: system gets the inputs of languages separately and the system identifies the language.

In this test case we gave input of four English language sentences for language identification, after the execution of the program system pop-ups the identification complete message and in the new window shows the input sentences with identified language and confidentiality of identification. We have calculated the time take by each sentence to identify the language. The average response time is taken by each sentence approximately 0.1s. The speed of execution may vary based on the ping of the internet.
Total time = (Average response time) x (No. of sentences) = (0. 1) x (4) = 0.4s
Case 2: system gets the inputs of languages separately and the system identifies the language.

In this test case we gave input of three Urdu language sentences for language identification, after the execution of the program system pop-ups the identification complete message and in the new window shows the input sentences with identified language and confidentiality of identification. We have calculated the time take by each sentence to identify the language. The average response time is taken by each sentence approximately 0.1s. The speed of execution may vary based on the ping of the internet.
Total time = (Average response time) x (No. of sentences) = (0. 1) x (3) = 0.3s
Case 3: system gets the inputs of languages separately and the system identifies the language.

| Item Name | Type | No. of Units | Per Unit Cost (in Rs) | Total (in Rs) |
|---|---|---|---|---|
| Total in (Rs) | 74000 | |||
| Report | Miscellaneous | 12 | 500 | 6000 |
| HDMI wire | Equipment | 2 | 2000 | 4000 |
| Hard Disk (For storing large dataset) | Equipment | 3 | 8000 | 24000 |
| Ram | Equipment | 1 | 8000 | 8000 |
| GPU invidia | Equipment | 2 | 16000 | 32000 |