节点文献
业界专家的媒体发言对公司股价影响的分析
【作者】 刘超;
【导师】 朱敏;
【作者基本信息】 上海师范大学 , 应用统计硕士(专业学位), 2016, 硕士
【副题名】基于文本挖掘的研究
【摘要】 随着互联网的高速发展,股民与网民的高度重叠,股票市场的信息结构也已经发生了深刻的变革。监管部门、上市公司、财经专家等不再是仅有的信息提供者,随着微博、微信、贴吧等自媒体的兴起,信息的发布早已是网民的基本权利。相对于普通网民,业界专家的媒体言论具有更高的关注度同时也有更高的可研究性。但是,如何充分利用网络信息一直是一个难题,原因在于网络信息具有海量、半结构化、时效性强的特点。本文通过网络爬虫抓取海量的原始文本,对抓取后的文本信息分词之后进行清洗处理,可以避免建模样本的随机性。之后对预处理后的文本数据分组,基于房产板块股市的涨跌情况,构建量化指标,进而分析业界专家的媒体言论对股市的影响。研究成果不仅对于判断股市价格变化有重要参考价值,而且对于外汇市场和期货市场等其他金融市场的价格变化也具有借鉴意义。本文不仅分析了专家影响力的不同对股市的影响程度,也分析了专家言论的某些词组出现的频率的不同与股市波动的关联性。最终,本文通过处理的海量数据集训练得到了朴素贝叶斯分类器,经过验证发现其预测能力具有较强的可信度。本文利用计算机科学中的信息检索技术和自然语言处理技术从海量的互联网新闻中挖掘出有用的市场信息,利用媒体资讯来弥补信息不对称,从而为股票预测分析提供了一条切实可行的路径。但由于本文建立的挖掘模型还不是很完善,模型缺乏连贯性,在之后的研究中,如果可以将预处理、分词、特征提取等挖掘模块集合起来,形成分析流程,对今后业界专家的媒体言论的信息挖掘更有帮助。
【Abstract】 With the rapid development of Internet, Shareholders are usually netizen, The information structure of the stock market has changed profoundly. Information release is already one of the basic rights of Internet users. Compared with the ordinary Internet users, The words of experts in the media has a higher attention. But, how to use the information is really a problem. Because the information from internet is huge and been changed all the time.In this paper, I use the network spiders crawling mass of the original text, separate the words from original text. After all I clean them to get useful words. Based on the real estate sector stocks rise and fall of the situation, I construct the quantitative indicators. Then analyzed the influence of the comments on the stock market. The research results not only has important reference value for judging the stock market price,but also be useful to analyzed of the foreign exchange market and futures market.This article not only analyses the influence degree of the experts with different influence on the stock market, also analyses different correlation between the stock market fluctuations and the frequency of the occurrence of some of the phrases from the expert. In the end, we got the Na?ve Bayesian Classifier based on the massive of data sets. Validated found its ability to predict is strong. In this paper, by using computer science information retrieval and natural language processing technology, I got useful information from a huge number of Internet.Use the media of information to compensate for the asymmetric information, and provides a feasible path for stock prediction analysis. But due to the mining model established in this paper is not very perfect, If we could get together with the preprocessing, segmentation, feature extraction, it will be more helpful for the mining of information from internet.
- 【网络出版投稿人】 上海师范大学 【网络出版年期】2017年 02期
- 【分类号】F832.51
- 【下载频次】246