Share EncyclopediaHome EncyclopediaCategories Switch Channel

Search engine classification and working principles

2026-07-26 00:081240NameNetworking

10, (c, d, e, f) points to the web page a with a link called "demon satan", or the more the source page (b, c, d, e, f) gives the link, the more relevant page a is to be seen in the user search for "demon satan" and the more the search engine works, the more three steps can be seen: extracting an indexed database from the internet and searching in the index database. 3. 1 the capture of web pages from the internet uses the spider system program, which automatically collects pages from the internet, automatically accesss the internet and climbs along all urls in any page to other pages, repeats the process and collects back all the pages that have crawled. 3. 2 establishment of an index database where the collected pages are analysed by the analytical indexing system program to extract relevant web-page information (including urls, code types, keywords contained in page contents, keyword location, time of generation)

11, larger links to other web pages, etc.) a large number of complex calculations based on a certain correlation algorithm that captures the relevance (or importance) of each page to each keyword in the content of the page and in the hyperlink, and uses these information to create a web-based index database. 3. 3 when the user enters a keyword search in the index database, the search system program finds all relevant pages in the web index database that match the keyword. Since all relevant web pages are already relevant to the keyword, the higher the correlation, the higher the ranking. Finally, the page generation system organises the links to the results and a summary of the contents of the page. Spider of the search engine usually re-accesss all web pages on a regular basis (with different cycles, possibly for days, weeks or months, or for networks of different importance)

Rationale for the directory search engine

12 the pages are updated with different frequency) and the web-indexed database is updated to reflect updates of web content, new web-page information is added, dead links are removed and reordered according to changes in web content and links. In this way, the content and changes of the web page are reflected in the results of the user query. Although the internet is only one, the capabilities and preferences of search engines vary, so the pages captured vary and the ranking algorithms vary. The databases of large search engines store several to several billion web-page indexes on the internet, amounting to thousands of gs or even tens of thousands of gs. However, even if the largest search engine created an indexed database of more than 2 billion pages, it accounted for less than 30 per cent of ordinary pages on the internet, and web data overlaps between different search engines were generally below 70 per cent. The important reason we use different search engines is because they can search for different content. On the internet

13, much more, the search engine can't access the index, and we can't search it. 3. 4 the rationale for extracting the pages is as follows: google, for example, is global, skynet is a source document for all of china. ... We can see a multiplicity of situations in a random web page (e. G. Through the browser's “see source file”) ... In addition to the text that we can normally see from the browser, there are a large number of HTML tags.... The size (bytes) of web-page-source files is usually about four times the size of their content, according to statistics

14 information not related to the main content. These situations present both challenges and new opportunities for effective information search, and we would simply point out that, in order to support the subsequent search service, some features that represent its content need to be extracted from the web-page source file. In terms of current awareness and practice, the key words are the most representative of this feature. So, as a basic task of the pre-processing phase, it is to extract the key words contained in the contents of the web-source document. For chinese, it is to remove the words from the text of the web page using a dictionary called "schematic software". After that, a web page is largely represented by a set of words, p. = t1, t2, tn. In general, we can get a lot of words, and the same word may appear repeatedly on a web page

Rationale for the directory search engine

In the middle of the expression, you have to remove words like "that" and "in" that have no meaning, which are referred to as "no word." thus, for a web page, the number of valid words is around 200. The elimination of duplicate or reproducing pages facilitates the reproduction, reproduction and re-issuance of web pages, and therefore we see a large number of duplications in web information.... According to statistics, the average repetition rate of web pages is about 4., i. E., when you see a web page through a url, there are on average three different urls that offer the same or essentially similar content. This is a positive phenomenon for a large number of internet users because of more access to information. But for the search engine, it's mostly negative; it's not only a consumer when collecting web pages

16, machine time and network bandwidth resources, and if they appear in the query results, which are meaninglessly draining the computer display resources, will also lead to user complaints that “so many repeats, give me one.” the elimination of duplication of content or subject matter is therefore an important task in the pre-processing phase. 3. 4. 3 analysis of links indicates that the large number of HTML tags creates some problems and new opportunities for pre-processing of web pages. From the point of view of information retrieval, if the system is dealing only with text of content, we can base ourselves on a “shared vocabulary assumption”, i. E. A collection of keywords included in the content, with the maximum word frequency (term flash or tf, tf) and the frequency of documents appearing in the document collection (d)Statistics like tf and df

17 it makes sense to indicate to some extent the relative importance of terms in a document or the relevance of certain elements. ... With HTML tags, the situation may be further improved, for example, where information is likely to be more important in the same document, and between than information between and between.... In particular, the link to other documents contained in HTML files is a subject of particular concern in recent years, which is seen as providing not only a relationship between web pages, but also an important role in judging the content of the web pages. For example, the words “herdings outside” are not available on the legendary home page, so that a search engine based solely on text analysis cannot return to the home page as a result

Rationale for the directory search engine

18. Querying the relevant list of results. The order of entries in the list is an important question. Because of the variety of users and the natural language style of the query, returning the same list to the same q0 is certainly not satisfactory to all q0 submitting users (or to the highest level of satisfaction). So the search engine actually pursues a statistical satisfaction. It is believed that google is now better than 100 degrees because in most cases the content returned by the former is better suited to the needs of the user than in all cases. There are many factors that need to be taken into account in order to sequence the search results, which will be discussed in depth. This is just an overview of what may be called “important” factors that may be developed during the pre-treatment phase. By definition, since they were formed during the pre-processing phase, they have nothing to do with the user query. How does one page matter more than another,

The core idea is that “many of the quotes are important.” the concept of “citation” is very well represented between web pages through the HTML superchain, as exemplified by the creation of pagerank as google's core technology. In addition, different features of web pages and literature have been noted, namely, that some pages are mainly large external links, with little specific thematic content per se, while others are linked to a large number of other pages. In a sense, this has created a relationship with a couple, which allows for the creation of another indicator of importance on the web. Some of these indicators can be calculated at the pre-processing stage or at the query stage, but all are part of the parameters used to sequence the results that eventually emerge during the query service phase. When we search different engines using the same keyword, the results vary. Differing results may also occur when searching for another keyword in the same search engine. This alternative learning has enabled me to learn a great deal about the search engine, improved my self-learning and hands-on skills, and made me aware of the many issues that require attention in searching for information, which are of great benefit to my future work. References: 1. 2. 34. 5. 6. 7.

Like 0
Report
Favorite 0
Tip 0
Comment 0
Share 0
MoreRelated Comments
No comments yet, be the first to comment