E-Commerce Websites are a great way of shopping online and having products delivered to our doorsteps. It has opened doors to countless opportunities for vendors/manufacturers to stretch their customer base globally. E-Commerce websites invite customers to purchase goods and products whe
Fake Review Detecting on E Commerce Website using Sentiment Analysis
E-Commerce Websites are a great way of shopping online and having products delivered to our doorsteps. It has opened doors to countless opportunities for vendors/manufacturers to stretch their customer base globally.
E-Commerce websites invite customers to purchase goods and products where they are only able to make decisions based on pictures provided by the sellers and slight description. The issue arises when malicious individuals use these platforms to benefit their personal interest by advertising appealing products which are otherwise fake and deceives the customer. Consumers all over the world suffer from this problem when shopping online which has a very negative impact on the E-Commerce websites. One way to avoid these sellers is to read the customer reviews that are posted by buyers who have formerly bought the product and used it. This is a solution but a tedious one as one product can have hundreds of thousands of reviews which leaves any customer inable read each and every review to judge or perform analytical comparisons.
The main problem is when deceptive product owners use these review features to benefit themselves. They produce fake reviews to either promote their own products and sometimes to demote competitors.
We aim to identify these reviews through the use of Natural Language Processing (NLP) with Sentiment Analysis (SA) also called opinion mining to detect fake reviews on the E-Commerce websites. NLP and SA are subfields of Artificial Intelligence (AI) that trains machine learning algorithms for specific purposes.
The main project objectives are:
The project has two important implementation steps.
For starters, a machine learning classification algorithm with the highest achievable accuracy will be opted and trained with the balanced dataset. For this purpose, we have 4 classes. Namely
We can multiple solutions for this problem. Initially, a Support Vector Machine (SVM) will be best suited as it draws a hyperplane for classification purpose. Secondly, Logistic Regression (LR) classification algorithm will be well suited as for the historical related works suggest LR achieving higher accuracy than other prominent classifiers.
For second, the higher performing classifier will be incorporated in the google chrome extension that in real time performs web scrapping from E-Commerce websites and filters out review section from the entire webpage. Applies NLP procedures and Sentiment Analysis. Which involves
1. Data acquisition
2. Data-preprocessing
3. Vectorization
4. Feature Extraction
5. Sentiment Analysis
6. Detection
7. Results collection
It then draws a trust-worthiness score for the customers using E-Commerce websites to see products purely on the basis of their review section. This score will be calculated on the ratio of Fake Review to Genuine Review.
The project will have following benefits.
Fake Review Detection - A Google Chrome Extension
The final deliverable will be a google chrome extension that is similar to an addon/plugin for Google Chrome Web Browser. It requires little to no user interference or interface whatsoever.
A machine learning classification algorithm will have already been trained for the purpose of review classification. At this point, the algorithm will be ready.
An array of raw-text will be passed as parameter to the method that has the classification algorithm in it. The array will be contain the plain text that is scrapped in real time from the webpage i.e. Amazon, Ebay.
The data is preprocessed to check for incosistencies namely blank spaces and null values. Stopwords are removed that can significantly affect feature extraction step from plain text for which we have used Term Frequency- Inverse Document Frequency (TF-IDF) that check for frequently occuring words.
Once the features are extracted, then it will be fed into the algorithm as an X-Label. Afterwards, the Y-Label is achieved that classifies plain text as one of the classes earlier discussed.
The results are collected for future effective highlighting of reviewers. In case, they are indulged in numerous products' review posting. It will mean they are most likely a spammer.
The process repeats for each and every review until the very last one. And a score is drawn resultantly.
| Item Name | Type | No. of Units | Per Unit Cost (in Rs) | Total (in Rs) |
|---|---|---|---|---|
| Geforce GTX 1660ti | Equipment | 1 | 65000 | 65000 |
| Total in (Rs) | 65000 |
There is need of automation of Stamp duty collection in board of revenue Sindh, because th...
The aim of the project is to design an intelligent device that is capable of monitoring th...
Cycling is a great form of exercise, it has many physical, mental, and social benefits. Ho...
Military Applications In air defense, it can be used for target detection, target recogni...
Suicidal ideations (SI), often called suicidal thoughts or ideas, is a broad term used to...