{"id":2897,"date":"2021-03-22T00:55:43","date_gmt":"2021-03-22T00:55:43","guid":{"rendered":"https:\/\/www.tejwin.com\/en\/insights\/numpy-pandas\/"},"modified":"2021-03-22T00:55:43","modified_gmt":"2021-03-22T00:55:43","slug":"numpy-pandas","status":"publish","type":"insight","link":"https:\/\/www.tejwin.com\/en\/insights\/numpy-pandas\/","title":{"rendered":"Numpy, Pandas"},"content":{"rendered":"<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" alt=\"\" class=\"wp-image-15479\" height=\"1024\" src=\"https:\/\/www.tejwin.com\/en\/wp-content\/uploads\/2026\/08\/image-215.png\" width=\"1024\"\/><\/figure>\n<p class=\"wp-block-paragraph\">Using NumPy and Pandas to start your first step of data analysis<\/p>\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\" id=\"ea0e\">After reading our previous articles, you might have already known how to get the data from TEJ API, store it into your computer, and update automatically! Then we are going to tell you\u00a0<strong>how to analyze this data by using these two important packages- Numpy and Pandas<\/strong><\/p>\n<\/blockquote>\n<h2 class=\"wp-block-heading\" id=\"9310\"><strong>Highlights of this article <\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>Numpy Intro\/Application<\/li>\n<li>Pandas Intro\/Application<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"4306\"><strong>Links related to this article<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>API Official Website:\u00a0<a aria-label=\"TEJ API Official Website (opens in a new tab)\" class=\"ek-link\" href=\"\/en\/about\/\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ API Official Website<\/a><\/li>\n<li>The Product Package:\u00a0<a aria-label=\"TEJ E SHOP (opens in a new tab)\" class=\"ek-link\" href=\"https:\/\/eshop.tej.com.tw\/E-Shop\/\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ E SHOP<\/a><\/li>\n<li>Source Code:\u00a0<a class=\"ek-link\" href=\"https:\/\/github.com\/tejtw\/TEJAPI_Python_Medium_DataAnalysis\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ GITHUB<\/a><\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\" id=\"9a65\">What is Numpy? How to use it?<\/h2>\n<p class=\"wp-block-paragraph\" id=\"a8b4\">Numpy is designed to conveniently and efficiently process n-dimensional and large-scale data arrays. With built-in functions, users could perform preliminary and rapid data processing.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Basic Application\uff0dSingle Dimension<\/strong><\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>import numpy as np\na = np.array([0, 0.5, 1.0, 1.5, 2.0]) #float ndarray -1-1.\nb = np.array(['a', 'b', 'c']) #string ndarray        -1-2.\nc = np.arange(0, 10, 2) #array([0, 2, 4, 6, 8])      -2.\nc[2:] #array([4, 6, 8])                              -3.\nc[:2] #array([0, 2])                                 -4.<\/code><\/pre>\n<p class=\"wp-block-paragraph\" id=\"c029\">Examples above:<\/p>\n<ol class=\"wp-block-list\">\n<li>Create a float data type array; string data type array<\/li>\n<li>Through np.arange() function, creating an array starts with 0, ends with 2, and the interval is 2.<\/li>\n<li>In python,\u00a0<strong>\u201d[]\u201d means select, and \u201d: \u201d means to\u2026.\u00a0<\/strong>But what we have to notice is that<strong>\u00a0the location of the first element is 0 instead of 1 in python.\u00a0<\/strong>Therefore<strong>, c[2:] means selecting the element from location 2 to the end (include the last element).<\/strong><\/li>\n<li>Same as above, but if we change from\u00a0<strong>c[2:] to c[:2]<\/strong>, which means selecting<strong>\u00a0<\/strong>elements<strong>\u00a0from start to location 1( location 2 is not included)!!<\/strong><\/li>\n<\/ol>\n<ul class=\"wp-block-list\">\n<li>Mathematical Tools<\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>a = np.arange(0, 30, 2) #array([0, 2, 4, ..., 28])\na.sum() #210                                      -1-1.\na.mean() #14.0                                    -1-2.\na.std() #8.640987                                 -1-3.\na.cumsum() #array([0, 2, 6, 12, ...,210])         -1-4.\nlst = [0, 2, 4]\nlst*2 = [0, 2, 4, 0, 2, 4]                        -2-1.\na+a #array([0, 4, 8, ..., 56])                    -2-2.\na*a #array([0, 4, 16, ..., 784])                  -2-3.<\/code><\/pre>\n<p class=\"wp-block-paragraph\" id=\"be3e\">Examples above:<\/p>\n<ol class=\"wp-block-list\">\n<li>the sum of array a; average; standard deviation; cumulative sum<\/li>\n<li>elements in array an add with the corresponding position; multiply with the corresponding position<\/li>\n<\/ol>\n<p class=\"wp-block-paragraph\" id=\"325a\">The first example is to use numpy built-in functions to calculate. In the second example, we can see the numpy vectorized computation. If we multiply a list(2\u20131) by 2,\u00a0<strong>the number of elements in the list will double instead of doubling the value.<\/strong>\u00a0But if it is numpy array(2\u20132, 2\u20133), it is possible to perform mathematical operations on the<strong>\u00a0corresponding positions of the elements<\/strong>\u00a0in the array~\ud83d\udcaa\ud83d\udcaa<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Basic Application\uff0dMultiple Dimensions<\/strong><\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>b = np.array([a, a*2]) #array([0, 2, 4, ..., 28],\n                              [0, 4, 8, ..., 56])\nb[0] #array([0, 2, 4, ..., 28])                   -1-1.\nb[0][1] #2                                        -1-2.\nb.sum(axis = 1) #array([210, 420])                -1-3.\nb.shape #(2, 15)                                  -2-1.\nb.reshape(5,6) #array([0, 2, 4, 6, 8, 10],        -2-2.\n                              ...\n                      [36, 40,   ..., 56]])<\/code><\/pre>\n<p class=\"wp-block-paragraph\" id=\"b192\">Examples above:<\/p>\n<ol class=\"wp-block-list\">\n<li>Select the first row of array b; select the second element of the first row of the array b; row sum of array b<\/li>\n<li>Shape(2*15) of array b; change to a new shape(2*15 -&gt; 5*6)<\/li>\n<\/ol>\n<p class=\"wp-block-paragraph\" id=\"6598\">Next, let\u2019s take a look at how numpy performs on multi-dimensional arrays. Similarly, we also use \u201c[]\u201d to select. The difference is that there are more elements that can be selected, so we can\u00a0<strong>use 2 \u201c[][]\u201d to select column and position respectively.<\/strong>\u00a0If we want to do some matrix operations, we can use\u00a0<strong>shape functions<\/strong>\u00a0in numpy to check and find the desired shape to do the calculation.~\ud83d\udcaa\ud83d\udcaa<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Other Applications\uff0dBoolean, Random Variables, Financial Functions<\/strong><\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>#boolean\nb &gt; 15    #array([False, False, ..., True],        -1-1.\n                 [False, False, ..., True])\nnp.where(b&gt;15, 1, 0) #array([0, 0, ..., 1],        -1-2.\n                            [0, 0, ..., 1])\n#random variable\nnp.random.normal(5, 2, 10)                         -2-1. \nnp.random.standard_normal(5)                       -2-2.\n#financial\npip install numpy_financial\nimport numpy_financial as npf\nnpf.fv(0.03, 5, 0, -1000) #1159.27                 -3-1.\n#fv(rate, nper, pmt, pv)                 \nnpf.irr([-95, 3, 3, 3, 103]) #0.0439               -3-2.\n#irr(values)<\/code><\/pre>\n<p class=\"wp-block-paragraph\" id=\"3976\">Examples above:<\/p>\n<ol class=\"wp-block-list\">\n<li><strong>Boolean:<br \/><\/strong>We can\u00a0<strong>directly use inequality<\/strong>(bigger than 15 in the example) to find the corresponding T\/F array\u00a0<strong>in numpy array<\/strong>\u00a0or use np.where() function to make a new way of judging T\/F (T is 1, F is 0 in the example).<\/li>\n<li><strong>Random Variables:<br \/><\/strong>Using different distributions in statistics to generate random variables, such as the\u00a0<strong>normal distribution<\/strong>\u00a0in the example(mean 5, std 2, 10 elements), and\u00a0<strong>standard normal distribution<\/strong>, and so on.<\/li>\n<li><strong>Financial Functions:<br \/><\/strong>In numpy, there is also a package designed for financial functions such as fv, pv, and irr which will be used when discounting. But we will need to install this package separately. All functions included in this package can be checked in\u00a0<a href=\"https:\/\/numpy.org\/doc\/1.17\/reference\/routines.financial.html\" rel=\"noreferrer noopener\" target=\"_blank\"><strong>HERE<\/strong><\/a>~.<\/li>\n<\/ol>\n<p class=\"wp-block-paragraph\" id=\"253b\">Numpy has many applications for data processing, so it is very difficult for us to tell you all of them in just one article\ud83d\ude22. Therefore, if you are interested in numpy, you can go through<strong>\u00a0<\/strong><a href=\"https:\/\/numpy.org\/doc\/stable\/\" rel=\"noreferrer noopener\" target=\"_blank\"><strong>Numpy Official Website<\/strong><\/a>\u00a0or leave the message below!\ud83d\udcaa\ud83d\udcaa<\/p>\n<h2 class=\"wp-block-heading\" id=\"8d7e\">What is Pandas? How to use it?<\/h2>\n<p class=\"wp-block-paragraph\" id=\"d7a1\">Pandas is a package that specializes in analyzing table data. Just like Excel, it presents data in a format we called DataFrame in order to help users analyze data more conveniently, especially for financial time series data.<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Basic Application<\/strong><\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>import pandas as pd\ndf = pd.DataFrame([1, 2, 3, 4],\n                  columns = ['Numbers'],\n                  index = ['index_a','index_b','index_c','index_d'])<\/code><\/pre>\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" alt=\"\" class=\"wp-image-15450\" height=\"164\" src=\"https:\/\/www.tejwin.com\/en\/wp-content\/uploads\/2026\/08\/image-212.png\" width=\"101\"\/><\/figure>\n<p class=\"wp-block-paragraph\" id=\"36b6\">From the codes above, we can<strong>\u00a0create a table with column name \u201cNumbers\u201d and row names \u201dindex_a, b, c, and d\u201d respectively.<\/strong><\/p>\n<pre class=\"wp-block-code\"><code>df.loc['index_a'] #Numbers 1                                  -1-1.\ndf.iloc[0:2] #refer to source code                            -1-2.\ndf * 2 #same as numpy                                         -1-3.\n#add \"Name\" column\ndf['Name'] = ['Amy', 'Bob', 'Catherine', 'Duke']              -2-1.\n#select whole column\ndf['Numbers']                                                 -2-2.\n#delete column\ndf.drop('column name', axis=1)                                -2-3.\n#Math\ndf['Numbers'].sum() #10                                       -3-1.\ndf['Numbers'].mean() #2.5                                     -3-2.\ndf['Numbers'].std() #1.291                                    -3-3.\n<\/code><\/pre>\n<p class=\"wp-block-paragraph\" id=\"dd7b\">Examples above:<\/p>\n<ol class=\"wp-block-list\">\n<li>Use loc and iloc to find the corresponding value. It should be noted that\u00a0<strong>loc is the name of the column\/row,\u00a0<\/strong>so we have to enter the name when selecting, while\u00a0<strong>iloc is the position corresponding to the element. For example(1\u20132), select the elements from the start to position 1 (2 is not included!).<\/strong><\/li>\n<li>Add; select; delete the column<\/li>\n<li>Sum of the whole df; average; standard deviation<\/li>\n<\/ol>\n<p class=\"wp-block-paragraph\" id=\"c73e\">Like the numpy arrays which we have mentioned earlier, in Pandas, we also use brackets\u00a0<strong>[\u201ccolumn name\u201d] to select or add columns.\u00a0<\/strong>But we will have to use the drop() function to delete columns. For operations, pandas dataFrame can perform basic statistical calculations in tables.<strong>~<\/strong>\ud83d\udcaa\ud83d\udcaa<\/p>\n<ul class=\"wp-block-list\">\n<li><strong>Basic Data Analysis<\/strong><\/li>\n<\/ul>\n<pre class=\"wp-block-code\"><code>import tejapi\ntejapi.ApiConfig.api_key = \u201c\u4f60\u7684api_key\u201d\ndf = tejapi.get('TWN\/EWPRCD', \ncoid = ['2330'],\nmdate={'gte':'2020-01-01', 'lte':'2020-12-31'}, \nopts={'columns': ['mdate','open_d','high_d','low_d','close_d']}, \npaginate=True\n)\n#Math\ndf.describe()\nnp.mean(df)\nnp.log(df)\n#Plot\ndf['close_d'].plot()<\/code><\/pre>\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" alt=\"\" class=\"wp-image-15452\" height=\"265\" src=\"https:\/\/www.tejwin.com\/en\/wp-content\/uploads\/2026\/08\/image-213.png\" width=\"314\"\/><figcaption class=\"wp-element-caption\">Descriptive Statistics Table<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\">The sample data we used for pandas data analysis is\u00a0<strong>2330.TW stock price daily data got from the TEJ API.\u00a0<\/strong>Then, most of the statistics that may be used further can be obtained through\u00a0<strong>describe() function<\/strong>(figure above\ud83d\udc46). If we want to do some operations on these values, we could\u00a0<strong>directly use numpy<\/strong>\u00a0to perform operations on the entire table!<\/p>\n<figure class=\"wp-block-image aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" alt=\"\" class=\"wp-image-15454\" height=\"328\" src=\"https:\/\/www.tejwin.com\/en\/wp-content\/uploads\/2026\/08\/image-214.png\" width=\"346\"\/><figcaption class=\"wp-element-caption\">Stock Price\uff08Daily\uff09<\/figcaption><\/figure>\n<p class=\"wp-block-paragraph\" id=\"c1da\">Last is the data visualization. There are several ways for users to plot the graph in python, and Pandas provides a very very easy one! If the chart we want to present\u00a0<strong>is not complicated\u00a0<\/strong>such as simple stock daily price, daily return, etc. We can<strong>\u00a0select the column and use the plot() function to directly see the result!\u00a0<\/strong>(figure above\ud83d\udc46)<\/p>\n<p class=\"wp-block-paragraph\" id=\"2f40\">The only thing we have to note here is that the\u00a0<strong>X and Y axes in the chart are the index and data you select respectively<\/strong>. That\u2019s why\u00a0<strong>we use a set_index() function to process our raw data at first.<\/strong><\/p>\n<h2 class=\"wp-block-heading\" id=\"5e52\">Conclusion<\/h2>\n<p class=\"wp-block-paragraph\" id=\"920a\">What we share with you this time is how to use Numpy and Pandas packages to do the data analysis. However, it is very difficult for us to explain all the functions included in these 2 packages. Therefore, if you have any question or interested in any topic, you could go to their websites or leave the message below \u2757\ufe0f\u2757\ufe0f Then, we will\u00a0<strong>go further into financial data analysis and applications in the next article<\/strong>, please look forward to it \u2757\ufe0f\u2757\ufe0f<\/p>\n<p class=\"wp-block-paragraph\" id=\"dae6\">Finally, if you like this topic, please click \ud83d\udc4f below, giving us more support and encouragement. Additionally, if you have any questions or suggestions, please leave a message or email us, we will try our best to reply to you.\ud83d\udc4d\ud83d\udc4d<\/p>\n<h2 class=\"wp-block-heading\" id=\"3bb0\">Links related to this article again!<\/h2>\n<ul class=\"wp-block-list\">\n<li>1\ufe0f\u20e3 API Official Website:\u00a0<a aria-label=\"TEJ API Official Website (opens in a new tab)\" class=\"ek-link\" href=\"\/en\/about\/\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ API Official Website<\/a><\/li>\n<li>2\ufe0f\u20e3 The Product Package:\u00a0<a aria-label=\"TEJ E SHOP (opens in a new tab)\" class=\"ek-link\" href=\"https:\/\/eshop.tej.com.tw\/E-Shop\/\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ E SHOP<\/a><\/li>\n<li>3\ufe0f\u20e3 Source Code:\u00a0<a class=\"ek-link\" href=\"https:\/\/github.com\/tejtw\/TEJAPI_Python_Medium_DataAnalysis\" rel=\"noreferrer noopener\" target=\"_blank\">TEJ GITHUB<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Using NumPy and Pandas to start your first step of data analysis After reading our previous articles, you might have already known how to get the data from TEJ API, store it into your computer, and update automatically! Then we are going to tell you\u00a0how to analyze this data by using these two important packages- [\u2026]<\/p>\n","protected":false},"featured_media":2896,"template":"","tags":[65,70],"insight_category":[16],"class_list":["post-2897","insight","type-insight","status-publish","has-post-thumbnail","hentry","tag-python","tag-tej-api","insight_category-quant-data-science"],"acf":[],"_links":{"self":[{"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/insight\/2897","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/insight"}],"about":[{"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/types\/insight"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/media\/2896"}],"wp:attachment":[{"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/media?parent=2897"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/tags?post=2897"},{"taxonomy":"insight_category","embeddable":true,"href":"https:\/\/www.tejwin.com\/en\/wp-json\/wp\/v2\/insight_category?post=2897"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}