{"id":476,"date":"2023-02-10T18:37:56","date_gmt":"2023-02-10T21:37:56","guid":{"rendered":"https:\/\/blackrat.pro\/blog\/?p=476"},"modified":"2023-02-10T10:59:24","modified_gmt":"2023-02-10T13:59:24","slug":"como-impedir-que-o-chatgpt-use-o-conteudo-do-seu-site","status":"publish","type":"post","link":"https:\/\/blackrat.pro\/blog\/como-impedir-que-o-chatgpt-use-o-conteudo-do-seu-site\/","title":{"rendered":"Como impedir que o ChatGPT use o conte\u00fado do seu site"},"content":{"rendered":"<p><span style=\"font-weight: 400\">H\u00e1 uma preocupa\u00e7\u00e3o com a falta de uma maneira f\u00e1cil de optar por n\u00e3o ter o conte\u00fado usado para treinar modelos de linguagem grandes (LLMs) como o ChatGPT. Existe uma maneira de fazer isso, mas n\u00e3o \u00e9 direto nem garantido que funcione.<\/span><\/p>\n<h2><span style=\"font-weight: 400\">Como as IAs aprendem com seu conte\u00fado<\/span><\/h2>\n<p><span style=\"font-weight: 400\">Os Large Language Models (LLMs) s\u00e3o treinados em dados origin\u00e1rios de v\u00e1rias fontes. Muitos desses conjuntos de dados s\u00e3o de c\u00f3digo aberto e s\u00e3o usados \u200b\u200blivremente para IAs de treinamento.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Algumas das fontes utilizadas s\u00e3o:<\/span><\/p>\n<ul>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Wikip\u00e9dia<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Registros do tribunal do governo<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">livros<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">E-mails<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Sites rastreados<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">Na verdade, existem portais, sites que oferecem conjuntos de dados, que fornecem grandes quantidades de informa\u00e7\u00f5es.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Um dos portais \u00e9 hospedado pela Amazon, oferecendo milhares de conjuntos de dados no Registro de Dados Abertos na AWS.<\/span><\/p>\n<p><span style=\"font-weight: 400\">O portal da Amazon com milhares de conjuntos de dados \u00e9 apenas um portal entre muitos outros que cont\u00eam mais conjuntos de dados.<\/span><\/p>\n<h2><span style=\"font-weight: 400\">Conjuntos de dados de conte\u00fado da Web<\/span><\/h2>\n<h3><span style=\"font-weight: 400\">OpenWebText<\/span><\/h3>\n<p><span style=\"font-weight: 400\">Um conjunto de dados popular de conte\u00fado da web \u00e9 chamado OpenWebText. O OpenWebText consiste em URLs encontradas em postagens do Reddit que tiveram pelo menos tr\u00eas votos positivos.<\/span><\/p>\n<p><span style=\"font-weight: 400\">A ideia \u00e9 que essas URLs sejam confi\u00e1veis \u200b\u200be contenham conte\u00fado de qualidade. No entanto, sabemos que, se o seu site estiver vinculado ao Reddit com pelo menos tr\u00eas votos positivos, h\u00e1 uma boa chance de que seu site esteja no conjunto de dados OpenWebText.<\/span><\/p>\n<h3><span style=\"font-weight: 400\">Rastreamento Comum<\/span><\/h3>\n<p><span style=\"font-weight: 400\">Um dos conjuntos de dados mais usados \u200b\u200bpara conte\u00fado da Internet \u00e9 oferecido por uma organiza\u00e7\u00e3o sem fins lucrativos chamada Common Crawl .<\/span><\/p>\n<p><span style=\"font-weight: 400\">Os dados do Common Crawl v\u00eam de um bot que rastreia toda a Internet. Os dados s\u00e3o baixados por organiza\u00e7\u00f5es que desejam usar os dados e, em seguida, limpos de sites com spam, etc. O nome do bot Common Crawl \u00e9 CCBot.<\/span><\/p>\n<p><span style=\"font-weight: 400\">O CCBot obedece ao protocolo robots.txt, ent\u00e3o \u00e9 poss\u00edvel bloquear o Common Crawl com Robots.txt e evitar que os dados do seu site entrem em outro conjunto de dados. No entanto, se seu site j\u00e1 foi rastreado, provavelmente j\u00e1 est\u00e1 inclu\u00eddo em v\u00e1rios conjuntos de dados.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Por\u00e9m, ao bloquear o Common Crawl, \u00e9 poss\u00edvel impedir que o conte\u00fado do seu site seja inclu\u00eddo em novos conjuntos de dados provenientes de dados mais recentes do Common Crawl.<\/span><\/p>\n<p><span style=\"font-weight: 400\">A string CCBot User-Agent \u00e9: CCBot\/2.0<\/span><\/p>\n<p><span style=\"font-weight: 400\">Adicione o seguinte ao seu arquivo robots.txt para bloquear o bot Common Crawl:<\/span><\/p>\n<ul>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">Agente do usu\u00e1rio: CCBot<\/span><\/li>\n<li style=\"font-weight: 400\"><span style=\"font-weight: 400\">N\u00e3o permitir: \/<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400\">Uma maneira adicional de confirmar se um agente de usu\u00e1rio CCBot \u00e9 leg\u00edtimo \u00e9 rastre\u00e1-lo a partir de endere\u00e7os IP da Amazon AWS. O CCBot tamb\u00e9m obedece \u00e0s diretivas da meta tag nofollow robots.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Use isso em sua meta tag de rob\u00f4s: &lt;meta name=&#8221;robots&#8221; content=&#8221;nofollow&#8221;&gt;<\/span><\/p>\n<h2><span style=\"font-weight: 400\">Bloqueando a IA de usar seu conte\u00fado<\/span><\/h2>\n<p><span style=\"font-weight: 400\">Os mecanismos de pesquisa permitem que os sites optem por n\u00e3o serem rastreados. Rastreamento comum tamb\u00e9m permite a desativa\u00e7\u00e3o. Mas atualmente n\u00e3o h\u00e1 como remover o conte\u00fado do site de conjuntos de dados existentes.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Al\u00e9m disso, os cientistas de pesquisa n\u00e3o parecem oferecer aos editores de sites uma maneira de optar por n\u00e3o serem rastreados.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Mat\u00e9ria completa:<\/span><a href=\"https:\/\/l.blackrat.pro\/867Ts\"><span style=\"font-weight: 400\">https:\/\/l.blackrat.pro\/867Ts<\/span><\/a><span style=\"font-weight: 400\">\u00a0<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>H\u00e1 uma preocupa\u00e7\u00e3o com a falta de uma maneira f\u00e1cil de optar por n\u00e3o ter o conte\u00fado usado para treinar modelos de linguagem grandes (LLMs) como o ChatGPT. Existe uma maneira de fazer isso, mas n\u00e3o \u00e9 direto nem garantido que funcione. Como as IAs aprendem com seu conte\u00fado Os Large Language Models (LLMs) s\u00e3o [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":477,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[88],"class_list":["post-476","post","type-post","status-publish","format-standard","has-post-thumbnail","category-how-to","tag-chatgpt"],"_links":{"self":[{"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/posts\/476","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/comments?post=476"}],"version-history":[{"count":1,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/posts\/476\/revisions"}],"predecessor-version":[{"id":478,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/posts\/476\/revisions\/478"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/media\/477"}],"wp:attachment":[{"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/media?parent=476"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/categories?post=476"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blackrat.pro\/blog\/wp-json\/wp\/v2\/tags?post=476"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}