Optimized Cross Domain Sentiment Classification through N-Gram Features and Machine Learning |
||||
|
|
||||
|
||||
BibTeX: |
||||
|
@article{IJIRSTV2I1039, |
||||
Abstract: |
||||
|
World Wide Web is full of blogs and forums in which provides the users a platform to share their views or remarks about diversified topics. Sentiment analysis refers to the use of natural language processing, text analysis and computational linguistics to identify and extract subjective information in source materials. It aims to determine the attitude of a speaker or a writer with respect to some topic or the overall contextual polarity of a document. The attitude may be his or her judgment or evaluation, affective state or the emotional state of the author when writing, or the intended emotional communication the author wishes to have on the reader. Sentiment classification aims to automatically predict sentiment polarity (e.g., positive or negative) of users publishing sentiment data (e.g., reviews, blogs). Although traditional classification algorithms can be used to train sentiment classifiers from manually labeled text data, the labeling work can be time-consuming and expensive. Meanwhile, users often use different words when they express sentiment in different domains. Words in different domain can be treated as positive or negative and vice-versa, e.g. hard knife and hard pillow. Cross-domain sentiment classification can be done by using a spectral feature alignment (SFA) algorithm to align domain-specific words from different domains into unified clusters, with the help of domain independent words as a bridge. The proposed work extends the technique proposed by S J Pan et.al. by including character level N gram features and shorthand internet notations, usually used in the web, into sentiment classification. An SVM based classifier is proposed to classify the polarity of the reviews. Domain Independence is achieved using cross domain words. Compared to previous approaches, this technique can classify the documents with much more accuracy as the shorthand notations are increasingly popular among the internet users. Extensive experiments are performed on real world datasets of twitter and it is demonstrate that inculcation of N gram features and shorthand notations can provide much better results for polarity classification within smaller false positive and false negative rates. |
||||
Keywords: |
||||
|
Sentiment Classification, Feature Extraction, Classification, Machine Learning etc. |
||||



