Online archives or archives of the online?

thumbnail_tendencias

At the end of 2020, we recommend some texts that put the future in perspective.

We highlight the theme of preserving online content presented in the ebook “Tendências 2021” (Trends 2021). The contribution of Daniel Gomes, the Arquivo.pt manager, was entitled “Arquivos online ou do online?” (Online archives or archives of the online?).

I was invited to write about the challenges and threats to online archives. The first question that came to me was what is meant by an “online archive”?

My concern lies in the “archives of the online” because there is not even an established awareness about their need, whether at an academic, governmental or individual level.

It is technologically impossible to preserve all information available online. But it is absurd not to be aware that we have to preserve some of the information online for short, medium and long term access.

The complete text (in Portuguese) is available at pages 23 to 26 of the open-access book “Tendências 2021”.

The challenge is to cultivate awareness about the importance of preserving content online by learning how to do it in practice.

Happy New Year!

Online Cafe with Arquivo.pt is back

Café com o Arquivo.pt

Last updated on August 23rd, 2022 at 04:17 pm

Café com o Arquivo.pt

Share this page: arquivo.pt/onlinecafe

Welcome to the second season of the Online Cafe with Arquivo.pt

Talk directly to the Arquivo.pt team and get answers to all your questions!  The Arquivo.pt launched a new cycle of team chats with you through online sessions. Brief introductory presentations will be given, leaving time to ask all your questions about how to get more out of Arquivo.pt or how to apply to the Arquivo.pt Awards.

Sessions held

21st session – Bilions of images to search on Arquivo.pt – all about the Arquivo.pt API

In March 2021 Arquivo.pt launched a new version with 1800 million images available. The search for images in web archives at this scale is unique in the world and innovative. The process used for indexing was explained in detail, in this session, as well as the best way to take advantage of these resources using the API to create new works based on images.

André Mourão, Ph.D. in Computer Science, is working on the indexing of the information, specially images from the Internet of the past to the present at Arquivo.pt (Portuguese Web Archive). He is a researcher working on ways to search and interpret multimedia data (e.g., images, text, video) effectively at large scales. He is also the co-creator of Revisionista.PT, uncovering post-publication edits in Portuguese news articles (with Flávio Martins) and an Associated Member at NOVA LINCS research center.

20th session – March 26 – The Online Centenary of the Great War

Daniela Major, invited speaker of the 20th Café with Arquivo.pt, presented a use case about historical research and old websites.  The commemoration of the Centenary of the 1st World War generated several strategies for international cooperation in the 21st century and such diversity is present in the information published on commemorative websites. In this session, Daniela Major will show how she used Web archives to start a study in the context of Contemporary History, as well as the methodological implications in her work.

Daniela Major is a PhD student in Digital Humanities at the School of Advanced Study, University of London. In 2019, within the scope of the ROSSIO Infrastructure, he started at Arquivo.pt a study on the celebrations of World War I based on the preserved contents of the Web. Currently, his PhD focuses on the media impact of the idea of Europe over the past few years. 15 years, in an effort to combine intellectual history with digital humanities.

This event, Cafe with Arquivo.pt, toke place for the first time 1 year ago, on the 27th march 2020. Cheers!

Special Session – March 2 – Arquivo.pt Award story and questions

Resume: The Arquivo.pt Award was created in 2018 in order to promote the use of Arquivo.pt for innotive works. How many candidates had applied since till now? What areas did the work focus on? What is the balance between Studies and Applications? Who were the winners and what were their main contributions? These and other questions were the guidelines for this session dedicated Award 2021.

This session was presented by Daniel Gomes and the Arquivo.pt team

18th session – Newspapers and web archives

This session was dedicated to the websites of Portuguese newspapers. Diogo Silva da Cunha talked about his first contact with Arquivo.pt and his approach, concretized in an “research route” as a way of delimiting the scope of his analysis. He also presented the results of his research about Correio da Manhã, Diário de Notícias, Expresso and Público newspapers.

Diogo Silva da Cunha is PhD student at the Institute of Social Sciences at the University of Lisbon, collaborator at the Center for Philosophy of Sciences at the same university. He recently published a study on the preservation of newspapers’ web pages in the book “O choque tecno-liberal, os media e o jornalismo: estudos críticos sobre a realidade Portuguesa”. In 2017, he participated in the Digital Humanities project “Investiga XXI” at Arquivo.pt.

17th session – january  15 – How to do an exhibition of old web pages without being an IT expert (tutorial)

This session present by Ricardo Basílio, the Arquivo.pt digital curator, is focused in practical aspects, when preparing an exhibition of old pages. As example: the use of long links, tipical of web archives, graphic aspects to be taken into account and navigation routes between webpages. The WordPress.com is used as platform to show how easy is build a web exhibition. The core aspects of dissimination content from Web archives have application in other platforms.

  • Video and presentation: (soon)

To know more about Portuguese newspapers in the Arquivo.pt see the exhibition Memória da Imprensa Portuguesa.

16th session – december 11 – Arquivo Económico .pt

Arquivo Económico .PT, authored by Nuno Bragança, 3rd place in Arquivo.pt Awards 2020, is a WebApp that allows discovering prices on web pages along time, over a set products in frequent use and compare them with current prices. Data are obtainned automatically from Arquivo.pt, processed and presented in an intuitive way for the common user. The possibility of comparing the present with the past based on information from the archived web shows how it can be useful not only to satisfy curiosity but also to support studies in many areas.

Query satisfaction of this presentation

15th session – november  24 – Extension Arquivo.pt

In this session we met the winners of 2nd place the Arquivo.pt Awards 2020. Rodrigo Marques and Hugo Silva talked about the their work “Arquivo.pt Extension” wich is a browser extension that allows users to search on Arquivo.pt. They showed through practical examples how the extension save time and helps the acess to the Arquivo.pt.

Query satisfaction of this presentation

Special session – World Digital Preservation Day 2020 – november 5

In November, World Digital Preservation Day is broadly celebrated and, to mark this international initiative, Arquivo.pt held an online session open to the community. The special guest of this session was the winner of the Arquivo.pt Award 2020, Miguel Ramalho, who told us about his work entitled “Desarquivo”.

Sessions held  between mars and july 2020

We improved the Arquivo.pt interface (Basileus release)

Thumbnail feature basileus version

Last updated on November 17th, 2020 at 07:00 pm

Arquivo.pt launched a new version, called Basileus, on November 11, 2020.

The purpose of this version was to improve the user experience when browsing through the different interfaces of Arquivo.pt.

Adjustments were made at the level of Web design which resulted in greater consistency in the structure of the code, in the graphic aspects and in the interactions, such as colors, fonts and buttons.

Print of release basileus of arquivo-pt replay Geocities

Figure 1: Search and replay interface of Web pages. In the figure, a replay of a Web page from Geocities historical collections on Arquivo.pt.

Help us to improve!

To help us, just search the Arquivo.pt using any device (e.g. laptop, mobile phone, tablet).

If you encounter any problems, please contact us!

Remember to always send the address of the page where you detected the problem.

To know more

World Digital Preservation Day 2020

WDPD2020-English-Portrait-RGB

Last updated on November 23rd, 2020 at 06:20 pm

WDPD2020-English-Landscape-RGB

On November 5, World Digital Preservation Day, Arquivo.pt held an online session open to the community.

Registration form (free but required)

The speaker for this session was the winner of the Arquivo.pt 2020 Award, Miguel Ramalho, who presented his work. “Desarquivo” is a web aplication that searches for entities on Arquivo.pt and return a graph.

As in 2017, 2018 e 2019, we invited everyone to get to know Arquivo.pt, and to use it in research and in the preservation of memory.

World Digital Preservation Day is promoted by the Digital Preservation Coalitium (UK) and an occasion for initiatives around the world, shared on social networks with the WDPD2020 hashtag.

Agenda:

November 5th

3:00 pm – Welcome! Presentation of the Arquivo.pt team (slides, 1 MB, PDF)
3:05 pm – Archive News – Daniel Gomes (slides, 2.6 MB, PDF)
3:15 pm – Desarquivo, 1st place in the Arquivo.pt Awards 2020, by Miguel Ramalho (slides, 3 MB, PDF)
3:45 pm – Questions
4:00 pm – Conclusion

Session video

Satisfaction query

Search the Geocities history!

thumbnail research_geocities

Last updated on September 23rd, 2021 at 03:30 pm

Geocities.com was the first major “social network” which enabled anyone to create their website and publish information on the Web. It was created in 1994, acquired by Yahoo in 1999 and shut down in 2009.

Initiatives have been emerging to preserve the content of Geocities, such as the Archive Team project which gathered 641 GB of information in 2009oOCities or Geocities.ws.

Arquivo.pt also integrated Geocities history in its collections!

Now, anyone can explore Geocities through the innovative tools provided by Arquivo.pt (e.g. full-text search, image search or API).

By making the historical collection of Geocities available, Arquivo.pt intends to contribute to the development of innovative studies in areas such as Arts, Humanities or Sociology (see a project summary).

Search Geocities now at: arquivo.pt/searchGeocities

Examples of Geocities preserved websites

Video Enhancing access to research the Geocities historical collection

Enhancing access to research the Geocities historical collection, Pedro Gomes, RESAW 2021 (slides)

 

Cross-lingual collection about the 2019 European Elections is available

print_europeanelections_q

Last updated on August 30th, 2022 at 10:46 am

Print European Elections 2019
Print from an archived page on Arquivo.pt: https://www.european-elections.eu

The special collection of web pages about the 2019 European Elections is available for search at Arquivo.pt.

To compile this collection, pages written in 24 European languages ​​were identified through automatic searches on the Bing search engine and suggestions from 17 European countries.

We emphasize the collaboration of the Publications Office of the European Union, which reviewed the list of search terms in the different languages ​​of the European Union.

Between May and July 2019, Arquivo.pt exhaustively collected pages related to the European Elections in several countries.

The resulting collection named “European Elections 2019” comprises 99 million web files that sum 4.8 Terabytes of information.

The technical report “A transnational crawl of the European Parliamentary Elections 2019 ” details the applied methodology. This methodology has been applied to generate other thematic collections such as about Covid-19.

We invited all citizens, especially the researchers, to try this service especially created to search the 2019 European Elections cross-lingual and international collection: https://arquivo.pt/ee2019

Video “A transnational and cross-lingual crawl of the European Parliamentary Elections 2019”

A transnational and cross-lingual crawl of the European Parliamentary Elections 2019, Ivo Branco, IIPC Web Archiving Conference and RESAW 2021 (slides)

To know more:

Collection about Covid-19 in Portugal

Thumbnail Covid-19 colletcion in Portugal

Last updated on June 18th, 2021 at 08:26 am

Banner Covid-19 colletcion in Portugal

Suggest web pages about Covid-19

Arquivo.pt invites everyone to suggest web pages that document the Covid-19 pandemic to be preserved for future access. Help us to keep a complete memory of the Portuguese live during this period.

Suggest pages using this form: https://tinyurl.com/arquivopt-covid19

Thousands of web pages to tell the story of the pandemic in Portugal

Arquivo.pt has been carrying out special collections of web pages related to the Covid-19 pandemic since March 2020.

“Future academics, scientists and journalists who are studying the Portuguese response to the Covid-19 pandemic will want to read first-hand testimonies of those affected, official records of the number of victims, and recommendations from doctors, politicians and scientists at the time” , Público newspaper, May 1, 2020 edition.

Daily, content was collected from a set of 106 sites on the theme of Covid-19. This set includes, for example, websites for the media, government, associations and university initiatives.

In another set are Twitter pages (108 identified in May), Youtube videos (815 identified in May) and also pages from Reddit and Git Hub.

Suggestions from the community were included. For example, Archivists from Sines (Portugal) collected local news related to Covid-19 (9 GB). The Revisionista.pt project also contributed and identified pages from newspapers. People sent suggestions through the public form.

Collaboration with IIPC for international collection

In February 2020, the International Internet Preservation Consortium (IIPC), the main organization on Web preservation, proposed to its members a collection about the Novel Coronavirus (Covid-19) outbreak.

Arquivo.pt contributed with 1 237 seeds, mainly in Portuguese. With successive contributions from other countries, the IIPC collection reached over 7 000 pages in July 2020.

A form is also available for anyone to suggest content for this international collection.

The IIPC collection “Novel Coronavirus (COVID-19)” is accessible via the Internet Archive Archive-it.

Arquivo.pt carried out 3 collections of the international collection compiled by the IIPC, the 1st on March 23 the 2nd on June 15 and the 3rd on late August, thus gathering international content useful for worldwide researchers.

Methodology for the selection of pages for the Covid-19 collection

We started by identifying terms related to the Coronavirus theme that included health, economic, political, geographic or organizational aspects.

Then, the Bing Azure service was used to automatically obtain, through a script, the following information for the first 10 results for each term: the page address, the title and the position in the results list.

Considering the list of results, it was decided which software would be used and which settings would be the best to collect the pages.

For example, in the case of a newspaper section dedicated to Covid-19, it was necessary to decide whether to record just one page or whether it makes sense to collect the entire site exhaustively.

Various types of software were used to collect the pages. For daily collections from 106 sites Heritrix was used. For capturing 108 Twitter accounts, Brozzler was chosen and for videos, manual capture using Webrecorder and Browsertrix.

Know more

Meet the winners of the Arquivo.pt Award 2020!

Card Meet the winners of the Arquivo.pt Award 2020

Last updated on October 2nd, 2023 at 10:18 am

The winners of the Arquivo.pt 2020 Award were announced by the Público newspaper, the official media partner of this year’s edition, which granted an honorable mention to the best work based on the contents of the newspaper. 29 candidate works were received.

The award ceremony toke place during Science 2020 – Meeting with Science and Technology, November 4, at the Lisbon Congress Center.

1st place – “Desarquivo”

The winner of the 10,000 euros prize was the work “ Desarquivo ” developed by Miguel Ramalho.

“Desarquivo” is a website that enables searching for named entities (e.g. people, organizations and places) and identify relationships among them, based on news published in online newspapers along time.

The search results are presented in the form of a graph or network of relationships that enables a journalist, researcher or any common citizen to dynamically explore the relationships among historical information preserved from the Web by Arquivo.pt.

For example, a user can explore ideological proximity among political parties along time.

2nd place – “Arquivo.pt Extension”

The 2nd prize in the amount of 3,000 euros was awarded to the work “ Extension Arquivo.pt ”,  a browser extension developed by Rodrigo Marques and Hugo Silva.

This extension enables users to perform advanced searches on Arquivo.pt directly from the browser , without having to leave the page they are currently viewing.

The “Arquivo.pt Extension” is available for download in the Chrome Web Store.

3rd place – “Arquivo Económico .pt”

The 3rd place winner received a prize of 2,000 euros and was awarded to the work “Arquivo Económico .pt” by Nuno Bragança.

The “Arquivo Económico .pt” organizes and presents information preserved by Arquivo.pt about the prices of products since the time of the Portuguese coin escudo.

As a result, we have a website that enables searching the price of consumer goods by different categories, such as supermarket, transportation or others, on given dates.

For example, users can easily know how much a trip from Lisbon-Porto or a cell phone call costed in 1999.

Honorable Mention granted by Público newspaper

Jornal Público, official partner of the 3rd edition of the Arquivo.pt Prize, awarded its Honorable Mention to the work “Jornal do Passado”, developed by Bruno Galhardo.

“Jornal do Passado” is a game for all ages, developed for Android, in which the users test their knowledge about news or events by guessing the date in which they occurred.

As a result, we have an app that enables searching the historical information preserved by Arquivo.pt in a pedagogical and fun way.

Image gallery

Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
20201104-EncontroCiencia-0140
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 no grande auditório do Centro de Congressos de Lisboa
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020
Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 20201104-EncontroCiencia-0140 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 no grande auditório do Centro de Congressos de Lisboa Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020 Entrega de prémios na sessão de encerramento do Encontro Ciência 2020

Replay with old browser and export results with the new version of Arquivo.pt

Exported results into an Excel sheet of a search for the word "universidade", university, limited to 10 items

Arquivo.pt launched a new version of its service on July 1, 2020 named Responsive.

The purpose of this version was to improve the user experience between different devices and add new features.

Replay a past webpage using a browser from the past

We added an option to view the archived page using a browser from the past. In the Options choose Replay with old browser and you will be redirected to the oldweb.today service that emulates browsers such as Netscape NavigatorMicrosoft Internet Explorer or NSCA Mosaic.

This external service is useful for research use cases, in areas such as Web design, Art, Communication or History,where it is necessary to access the original visual aspect of a page from the past in the most reliable way possible.

Web page of the European Union in 1996 using the Oldweb.Today service
Web page of the European Map of WWW/NIR sites in 1996 using the Oldweb.Today service

Try this new option from Arquivo.pt to replay the European Map of WWW/NIR sites in 1996 using a contemporary browser or any other historical page using the Oldweb.Today service.

You may have to wait a while for your request to be processed but it is always faster than having to install a browser from the past on your computer.

Export search results to spreadsheet format

This new function enables users to save their search results for further treatment and analysis. This is specially useful to perform thorough research about a given topic.

After a search, in the Options, just choose one of the available formats to export the obtained results: XLSX, CSV or TXT.

Exported results into an Excel sheet of a search for the word "universidade", university, limited to 10 items
Exported results into an Excel sheet of a search for the word “universidade”, university, limited to 10 items

More on the Responsive milestone

New version of Arquivo.pt (Webapp release)

Webapp release on mobile version example

Last updated on October 12th, 2020 at 11:52 am

Arquivo.pt launched a new version of its service on April 15, 2020 named WebApp.

The purpose of this version was to standardize the user experience between different devices and reduce maintenance costs by removing components with redundant functions.

Its main novelty is the combination of the desktop and mobile interfaces in a single user interface.

The old desktop version has been disabled and the mobile version has evolved to work on various types of devices and screen sizes.

Webapp release desktop and mobile

New design of the homepage

 

Try the new image and page search

Webapp release search in english

New user interfaces for image or text search

Help us to improve!

To help us, just search the Arquivo.pt using any device (e.g. laptop, mobile phone, tablet).

If you encounter any problems, please contact us!

Remember to always send the address of the page where you detected the problem.

More information